First Time Implementing RAG

I built Repo Recall because I kept forgetting how my own projects worked. A few months after building something, I would look at my old code and have no idea what I had written. So I built a tool that lets me load any GitHub repo, select files, and chat with them in context.
You could ask: VS Code already has AI, so why build this? You can. This one runs in the browser. I do not have to clone the repo or open an editor. I open the site, load the repo, and ask.
Stack: React (Vite), Express, Postgres with pgvector, Gemini for embeddings and chat
Deploy: frontend on Vercel, API on Render
The Architecture
You paste a GitHub URL, pick a branch, and select the files you care about. Those paths go to the server. The server still needs the actual text before it can embed anything. That was the first thing that slowed me down.
My first plan was simple. Every file has a GitHub URL, so I would call GitHub, get the content, embed it, and move on. It works for two files. For a real repo it means one network call per file, and that gets slow fast.
I spent time I could have spent sleeping, and then I found adm-zip.
I download the whole branch as one zip. That download is raw bytes. I turn those bytes into a Buffer and hand the buffer to adm-zip. After that I can ask it, "give me the content of this path," and it reads that file out of the zip. The repo is already on the server. I do not call GitHub again for every file.
Then the next problem showed up. Say I already embedded a file. The user unchecks it, then checks it again later. Do I embed it a second time? And if the file itself has not changed, embedding it again is just wasted work. Here is how i designed this:
First time you select a file, there is no saved hash, so you embed it and store the chunks.
Same hash means the file has not changed, so you skip embedding.
Different hash means the content changed, so you delete the old chunks and embed again.
Deselected files stay in the database. You only set
activetofalse, so chat ignores them.
The hash is the embedding model plus the file content.
What I Learned?
RAG is me deciding what the model is allowed to see. It does not get the whole repo. It gets a few chunks I picked. While reading IBM's RAG docs I fell into generative AI, NLP, and neural networks. I only read enough to get curious. I did not use all of that in this project.
The part I actually used is the vector database.
An embedding model takes a chunk of code and returns a list of numbers. I don't choose those numbers. The model does, and the length of the list is fixed by the model. I store that list in Postgres with pgvector, next to the chunk it came from.
When I ask a question, that question becomes another list of numbers, from the same kind of model. Then I compare the two lists with cosine similarity. It is not "are these numbers equal." It is "do they point the same way." Closer means the chunk is more like the question. I keep the closest ones (cosine distance <0.7 ) and put only those in the prompt. That is the constraint.
I also learned that a
UNIQUEconstraint builds an index. Onindexed_files,UNIQUE (repo_id, path)is how Postgres finds "have I embedded this file before" without scanning the whole table.And a zip is not a folder I unpacked. At the last of the file is a directory of every path. The file contents stay compressed inside.
adm-ziponly decompresses a file when I ask for that one path. TheBufferis the zip itself, not the source code.
