Skip to content

Retrieval lab

Ask a question about my work and watch how a RAG system would find the context: keyword search, embedding search, the fusion of both, and a graph step that pulls in related work through shared entities.

Hybrid retrieval over this portfolio

150 chunks ยท Lexical only
[1]Veriflowoverview
lex #1rrf 16.4

Veriflow takes an investigation objective, such as finding which B2B SaaS customers are at risk of churning and why, and returns a report in which every claim is tied to evidence and checked before it is shown. Any customer-facing action it proposes waits for a human approval.

[2]Veriflowresult
lex #2rrf 16.1

Claims supported after verification: 81.8% (5-account run, 33 claims; 0 rejected claims reached the report).

[3]Veriflowfeature
lex #3rrf 15.9

Three-stage claim verification with evidence links on every claim in the final report.

[4]Veriflowproblem
lex #4rrf 15.6

LLM agents are good at producing fluent conclusions and bad at telling you which of those conclusions are supported. In an account-management setting, an unsupported claim that reaches a customer is worse than no claim at all.

[5]Veriflowsummary
lex #5rrf 15.4

Veriflow. An agent platform that has to prove every claim it makes before a human signs off. Built with FastAPI, Pydantic v2, SQLAlchemy 2, PostgreSQL 16, pgvector, Redis, Celery, Next.js, TypeScript, OpenTelemetry, Docker Compose.

Graph expansion

No other item shares two or more uncommon entities with the retrieved sources.

Items not retrieved directly but linked to the results through at least two shared technologies or categories. Shared entities are weighted by rarity, and ones attached to more than a quarter of items (like Python) are ignored.

Context window
Chunks
5
Approx. tokens
296
Answer using only the context below. Cite sources by number.

Question: How are LLM claims checked before a report?

Context:
[1] (Veriflow, overview)
Veriflow takes an investigation objective, such as finding which B2B SaaS customers are at risk of churning and why, and returns a report in which every claim is tied to evidence and checked before it is shown. Any customer-facing action it proposes waits for a human approval.

[2] (Veriflow, result)
Claims supported after verification: 81.8% (5-account run, 33 claims; 0 rejected claims reached the report).

[3] (Veriflow, feature)
Three-stage claim verification with evidence links on every claim in the final report.

[4] (Veriflow, problem)
LLM agents are good at producing fluent conclusions and bad at telling you which of those conclusions are supported. In an account-management setting, an unsupported claim that reaches a customer is worse than no claim at all.

[5] (Veriflow, summary)
Veriflow. An agent platform that has to prove every claim it makes before a human signs off. Built with FastAPI, Pydantic v2, SQLAlchemy 2, PostgreSQL 16, pgvector, Redis, Celery, Next.js, TypeScript, OpenTelemetry, Docker Compose.
Everything runs in your browser. BM25 is computed locally; dense search uses all-MiniLM-L6-v2 (quantized ONNX, about 23 MB, downloaded from Hugging Face only when you load it). Ranks are fused with weighted reciprocal-rank fusion (k = 60). No language model is called: the right-hand panel shows the context a RAG system would send to one.

Algorithm demos

Each runs in your browser and reimplements a repository's notebook or C++ code. Where the notebook recorded results, the demo is tested against them.