← All work

AI Lab · LangGraph

Agentic workflows & RAG experiments

Our internal lab for stateful, multi-step LangGraph agents and retrieval-augmented generation pipelines — the proving ground where AI patterns get battle-tested before they go into client products.

LangGraph agentsRAG pipelinesStateful workflowsProduction-proven
Type
Internal R&D
Focus
Agentic AI · RAG · Evals
Stack
LangGraph · Python · pgvector
Scope
Ongoing

The challenge

Where production AI patterns get proven before they ship

We don't put AI patterns into client products without proving them first. The lab is where we build and break stateful multi-step agents, tune RAG pipelines with real evals, and figure out where rules-first beats LLM-first. The patterns that survive become the AI inside products like Kova's match scoring and USDT's field mapper.

What we delivered

  • Production-proven AI patterns before client deployment
  • RAG with grounding evals, not vibes
  • Rules-first architecture validated for deterministic use cases
  • On-device AI patterns proven before production use
  • Multi-agent orchestration battle-tested
  • LangSmith tracing for full observability

Features

What we built

Six core capabilities that make the product what it is.

Stateful LangGraph Agents

Multi-step agents with persistent state, conditional branching, and human-in-the-loop interrupts — validated in the lab before client production.

RAG Pipeline Tuning

Retrieval pipelines with pgvector + hybrid BM25/dense search, chunk strategy experiments, and answer grounding evals — not intuition.

Rules-first vs LLM-first Research

Systematic experiments on where deterministic rules outperform LLMs (mapping, scoring) vs where LLMs genuinely add value (interpretation, generation).

Eval Harnesses

Automated evaluation pipelines for accuracy, groundedness, and latency across agent versions. A pattern must pass evals before it ships.

On-device Model Experiments

Testing TensorFlow.js and ONNX models for on-device inference — validating patterns like FeminineXP's AR mirror before production commitment.

Multi-agent Patterns

Supervisor/worker agent architectures, parallelism, and error recovery — the foundation patterns for production AI orchestration.

Stack

Built with the right tools

No dogma about tools. We assessed what the product needed, then chose the stack that delivers fastest and stays maintainable long-term.

LangGraphOpenAIPythonpgvectorPostgreSQLTensorFlow.jsONNX RuntimeLangSmith

Ready when you are

Build something like this?

Tell us what you're working on. We'll map the right approach and stack for your context.