Semantic Search & Graph Retrieval
View Repository SourceAn embedded property graph and retrieval pipeline indexing 141 sessions of AI coding history with BM25 FTS and semantic search.
Searching through transcripts of AI agent sessions is fundamentally different from searching standard documentation. Agent sessions are mostly machine output, not conversation, and naive transcript searches quickly drown in directory listings and code dumps.
This project is an indexing and retrieval pipeline built over my own AI coding-agent history: 141 sessions, 8,731 conversational blocks, 17,296 chunks, and 156 distinct tools across 45 projects.
Intelligent Discrimination and Chunking
To make the data searchable, I implemented a content_kind discriminator that separates human prose from tool invocations and results. A default query searches the 4,755 prose chunks, while error-hunting queries can target specific tool outputs deliberately.
Turns are sub-chunked to ~250 tokens with sequential linkage. This prevents embedding models from silently truncating oversized inputs—a subtle defect that degrades retrieval quality invisibly.
Traversable Graph Architecture
Instead of a flat index, sessions are modeled as a traversable graph. Projects, chronological sessions, chunks, and tool invocations act as first-class nodes. Retrieval can be rigorously scoped by project, working directory, Git branch, or even which specific tool a turn called.
Practical Utility: Credential Auditing
A searchable archive of agent transcripts is also a searchable archive of everything those transcripts leaked. I wrote a credential auditor over this corpus that scanned 269 files and surfaced 144 potential exposures across eight secret classes, allowing me to plug security gaps proactively.