The AI Assistant on this site is not a wrapper around a chatbot API with a system prompt that says "you are Vishnu, answer about his work." It is a real retrieval-augmented generation pipeline over this site's own content, built the same way I would build one for a client.

Chunking the site, not just the courses

Every page that carries real information, profile, skills, experience, education, courses, articles, consultancy copy, case studies, gets broken into chunks and re-indexed whenever the underlying content changes. Each chunk keeps its own source type and a link back to the page it came from, so an answer can cite exactly where it got its evidence instead of generating a vague summary.

Embeddings and retrieval

Every chunk is embedded with an NVIDIA NIM model and stored in Postgres with pgvector. A question gets embedded the same way, then matched by cosine distance against every stored chunk. No separate vector database, Postgres already holds the content, so it holds the vectors next to it. That is one less moving part to operate and one less place for the data to drift out of sync.

A relevance gate before the LLM is ever called

Not every question deserves an LLM call. Before generating anything, the pipeline checks whether the closest retrieved chunk is actually a good match for the question. A question that is not actually about my work, general trivia, something unrelated, gets a fixed answer immediately, without ever reaching the LLM. That keeps the bot on topic and avoids paying generation cost on questions it has nothing grounded to say about.

Retrieval respects who is asking

Some courses on this site are paid, and their lesson content is gated behind a real purchase. The chatbot indexes that gated content too, because a buyer should be able to ask it questions while learning. But retrieval filters every gated chunk against the asking visitor's own unlocked course IDs. A visitor who has not bought that course simply never gets that chunk back, so the bot cannot become a free side-channel around a paywall.

Caching answers, not just embeddings

Repeat questions are common, and a full retrieval-plus-generation round trip is the most expensive part of the pipeline. A semantic answer cache sits in front of it: if a new question is nearly identical to one already answered, the cached answer is served instantly instead of calling the LLM again. The match threshold for a cache hit is deliberately tight, far tighter than the threshold used for ordinary retrieval, because a wrong cache hit is worse than a slow answer. Cached rows that cite gated course content stay scoped to that course's own buyers, the same access check retrieval itself already applies.

Why bother with all of this for a portfolio site

Because a RAG pipeline that only works in a slide deck is not a RAG pipeline. The version that actually matters has to handle paywalls, repeat traffic, and off-topic questions gracefully, the same problems a client-facing deployment hits on day one. Building it on my own content, where I can see every answer and every retrieved chunk, is how I know the pattern actually holds before I would recommend it to anyone else.

If you want to see it in action, ask the AI Assistant on this site a question about my work. The answer it gives you came from exactly the pipeline described above.