LAB_01RunningLive experiment

Agent orchestration patterns

Hypothesis
For multi-step tasks with clear sub-goals, an explicit graph is more predictable than a free-form supervisor, at a lower token cost.
Experiment
The same task set is run through graph, router and supervisor topologies with identical tools and models.
Finding
In progress. No conclusion published yet.
Measuring
Task completion, steps per task, token spend, recovery after a failed tool call.

LangGraph · CrewAI · custom task harness

LAB_02In useBenchmark

RAG evaluation harness

Hypothesis
Groundedness and citation precision can be scored mechanically against a golden dataset, without leaning on an LLM judge for every answer.
Experiment
Answers are checked claim by claim against the cited spans in a golden set of question–answer–citation triples.
Finding
Used as the regression gate for LexAI; the harness itself is still being refined.
Measuring
Groundedness, citation precision, refusal rate on out-of-corpus questions.

Qdrant · Ragas · golden triples

LAB_03RunningLive experiment

LLM routing by cost curve

Hypothesis
Many requests do not need a frontier model; routing by estimated difficulty can cut cost without a visible drop in quality.
Experiment
Each request is classified by difficulty and routed between a small and a frontier model under a fixed latency budget.
Finding
In progress. Collecting a large enough sample to compare.
Measuring
Cost per request, latency, quality parity against frontier-only answers.

LiteLLM · custom telemetry · Prometheus

LAB_04In useTooling

AI observability traces

Hypothesis
Agent failures are only debuggable when every tool call, retry and token is visible as a span in one trace.
Experiment
OpenTelemetry instrumentation around agent steps, tool I/O and model calls, viewed as a single trace per request.
Finding
Adopted as the default way to debug agent runs.
Measuring
Coverage of agent steps by spans; overhead added to each request.

OpenTelemetry · Jaeger

LAB_05Collecting dataBenchmark

Structured-output reliability

Hypothesis
Schema adherence degrades in predictable ways under truncated, noisy and adversarial inputs, and those failures can be recovered.
Experiment
Frontier and open-weight models are given the same schemas under clean and degraded inputs; failures are classified and retried.
Finding
In progress. Failure taxonomy being built.
Measuring
Valid-schema rate, failure type, success rate of validation-and-retry.

Pydantic · Instructor

LAB_06DraftingResearch note

Memory compaction for agents

Hypothesis
An agent can drop most of its conversation history without losing quality, if what it keeps is chosen well.
Experiment
Recursive summarisation, sliding windows and selective state retention are compared on long-running tasks.
Finding
Open question. Notes in progress.
Measuring
Task quality versus tokens retained.

Vector store · evaluation matrix

Next step

Have a system that needs
to hold up in production?