RAG & Agentic LLM Pipeline
A modular RAG pipeline over YouTube transcripts, extended into semantic search, a ReAct agent and an MCP server, packaged as a tested Python module.
- Status
- shipped
- Date
- Stack
- Python
- LangChain
- OpenAI API
- Pinecone
- ChromaDB
- Hugging Face
- pytest
Problem
I wanted to ask questions of long video content and get answers grounded in what was actually said, not what a model guessed. That meant a pipeline that could ingest transcripts, find the right passages and cite them, built as a module I could test and reuse rather than a notebook I would throw away.
- Role
- Sole engineer (DataCamp certification project)
Architecture
Chunking
YouTube transcripts split into chunks sized for retrieval
Embeddings
Each chunk embedded through the OpenAI API into namespaced Pinecone indexes
Retrieval
The question embedded the same way; the nearest chunks come back as the only context
Evals
pytest checks each seam, from ingest to answer, so a retrieval change cannot regress unseen
What was built
A tested Python module (pytest, PEP 8) that ingests, embeds, retrieves and generates.
Ingestion
YouTube transcripts are ingested and split into chunks sized for retrieval.
Embedding and retrieval
Each chunk is embedded through the OpenAI API and stored in namespaced Pinecone indexes. At question time, the question is embedded the same way and the nearest chunks come back as context, so answers stay grounded in the transcript.
Semantic search, agent and MCP
I reused the same techniques for a ChromaDB-based semantic search module, then built a LangChain agent that follows the ReAct pattern and calls custom tools. A minimal MCP server exposes those tools so an external LLM can use them too.
This project was the capstone of DataCamp’s 10-course, 29-hour Associate AI Engineer certification, which covered prompting, embeddings, RAG vs. fine-tuning trade-offs and LLMOps fundamentals.
Results
- Ingests, chunks and embeds YouTube transcripts, then answers from retrieved context only
- Retrieval runs through namespaced Pinecone indexes
- The same retrieval core powers a ChromaDB semantic search module, a LangChain ReAct agent and a minimal MCP server
- Packaged as a Python module with pytest tests and PEP 8 style, not a one-off script
Lessons
- Retrieval quality is decided at chunking time; the prompt cannot recover context that was never retrieved
- Writing tests for a pipeline forces clear seams between ingest, embed, retrieve and generate
- Exposing tools over MCP made the agent's tools usable by any LLM client, not just my own code
Next ProjectBusiness Model Canvas Builder →