← Work

RAG & Agentic LLM Pipeline

A modular RAG pipeline over YouTube transcripts, extended into semantic search, a ReAct agent and an MCP server, packaged as a tested Python module.

Status
shipped
Date
Stack
  • Python
  • LangChain
  • OpenAI API
  • Pinecone
  • ChromaDB
  • Hugging Face
  • pytest

Problem

I wanted to ask questions of long video content and get answers grounded in what was actually said, not what a model guessed. That meant a pipeline that could ingest transcripts, find the right passages and cite them, built as a module I could test and reuse rather than a notebook I would throw away.

Role
Sole engineer (DataCamp certification project)

Architecture

TranscriptYouTube Chunkretrieval-sized EmbedOpenAI API Pineconenamespaces Question Retrievenearest chunks Answerby the LLM ChromaDBsemantic search ReAct agentLangChain MCP serverfor any LLM
  1. Chunking

    YouTube transcripts split into chunks sized for retrieval

  2. Embeddings

    Each chunk embedded through the OpenAI API into namespaced Pinecone indexes

  3. Retrieval

    The question embedded the same way; the nearest chunks come back as the only context

  4. Evals

    pytest checks each seam, from ingest to answer, so a retrieval change cannot regress unseen

What was built

A tested Python module (pytest, PEP 8) that ingests, embeds, retrieves and generates.

Ingestion

YouTube transcripts are ingested and split into chunks sized for retrieval.

Embedding and retrieval

Each chunk is embedded through the OpenAI API and stored in namespaced Pinecone indexes. At question time, the question is embedded the same way and the nearest chunks come back as context, so answers stay grounded in the transcript.

Semantic search, agent and MCP

I reused the same techniques for a ChromaDB-based semantic search module, then built a LangChain agent that follows the ReAct pattern and calls custom tools. A minimal MCP server exposes those tools so an external LLM can use them too.

This project was the capstone of DataCamp’s 10-course, 29-hour Associate AI Engineer certification, which covered prompting, embeddings, RAG vs. fine-tuning trade-offs and LLMOps fundamentals.

Results

  • Ingests, chunks and embeds YouTube transcripts, then answers from retrieved context only
  • Retrieval runs through namespaced Pinecone indexes
  • The same retrieval core powers a ChromaDB semantic search module, a LangChain ReAct agent and a minimal MCP server
  • Packaged as a Python module with pytest tests and PEP 8 style, not a one-off script

Lessons

  • Retrieval quality is decided at chunking time; the prompt cannot recover context that was never retrieved
  • Writing tests for a pipeline forces clear seams between ingest, embed, retrieve and generate
  • Exposing tools over MCP made the agent's tools usable by any LLM client, not just my own code