An agent platform that has to prove every claim it makes before a human signs off.
Complete backend and frontend with a documented test suite; runs locally with Docker Compose. No public deployment.
Overview
Veriflow takes an investigation objective, such as finding which B2B SaaS customers are at risk of churning and why, and returns a report in which every claim is tied to evidence and checked before it is shown. Any customer-facing action it proposes waits for a human approval.
Problem
LLM agents are good at producing fluent conclusions and bad at telling you which of those conclusions are supported. In an account-management setting, an unsupported claim that reaches a customer is worse than no claim at all.
Approach
A planner turns the objective into a task graph. Specialised agents gather usage data, retrieve documents with customer-scoped pgvector search, and score risk. Every claim the investigator writes then goes through three verification stages: ID validation, lexical rules, and a model verdict that the rules can cap. Only supported claims reach the report, and the single write tool in the registry creates an approval record instead of acting directly.
Architecture
Select a component to see what it does. Blue packets show the direction data moves.
- Objective to Planner
- Planner to Data agent (tasks)
- Planner to Retrieval
- Planner to Risk agent
- Data agent to Investigator
- Retrieval to Investigator (evidence)
- Risk agent to Investigator
- Investigator to Verifier (claims)
- Verifier to Report (supported)
- Report to Human approval (actions)
What it does
- Planner that compiles an objective into a dependency graph of tasks, executed by Celery workers.
- Seven agents: four LLM-backed (Planner, Investigator, Verifier, Reporter) and three deterministic (Data, Risk, Retrieval).
- Customer-scoped retrieval over pgvector with an HNSW cosine index.
- Three-stage claim verification with evidence links on every claim in the final report.
- Twelve-tool registry with permissions; the only write tool proposes an action for human approval.
- Evaluation CLI covering ranking quality, retrieval scope, and verification outcomes.
- Next.js operator console: runs, live task events, reports, customers, approvals, evaluation, and system health.
Technical challenges
- Keeping a model verdict from overriding hard evidence: the rule stages cap what the model is allowed to mark as supported.
- Making runs observable while they execute, which meant persisting task events rather than only final state.
- Running the full system without paid keys: a deterministic fake LLM provider and a local hashed embedder keep the pipeline testable end to end.
Engineering decisions
- An OpenAI-compatible HTTP client instead of a vendor SDK, so any compatible endpoint can be swapped in.
- Deterministic agents wherever a model is not needed (data access, risk scoring, retrieval) to keep cost and variance down.
- Synthetic data from 12 customer personas generated from templates, so the dataset is reproducible and contains no real customer information.
Recorded results
- Risk ranking ROC-AUC
- 0.841
- Synthetic 3,000-customer dataset
- Source: AgenticAIPlatform/README.md, Evaluation results
- Precision at 25
- 0.96
- Top 25 accounts by risk score
- Source: AgenticAIPlatform/README.md
- Claims supported after verification
- 81.8%
- 5-account run, 33 claims; 0 rejected claims reached the report
- Source: AgenticAIPlatform/README.md
- Cost of that run
- $0.0074
- 8 LLM calls, 26,553 tokens
- Source: AgenticAIPlatform/README.md
Limitations
- Evaluated on synthetic data only.
- Single tenant, with development-level API-key auth.
- The default local embedder is a hashed bag-of-features vector, not a semantic model; semantic retrieval needs an embedding API key.