Skip to main content
AI Python Observability
View as Markdown Suggest changes

Langfuse in Production: From Failing Evals to Targeted Fixes

· Reading time: 6 min
Langfuse in Production: From Failing Evals to Targeted Fixes

In a recent AI-powered support project, Langfuse connected production monitoring, evaluation history, and a Cursor-based triage loop. What mattered was keeping the full trace of every failed eval case, tied to a commit, so we could tell whether to fix the eval data, a mock, a prompt or the code.

1. Monitoring production and evaluation flows

We treat observability as a port: the app depends on a MonitoringPort interface, and the Langfuse implementation lives in an adapter. That keeps the core logic free of SDK details and makes it easy to test with a no-op or mock.

The adapter exposes an observe decorator that wraps any function and sends a span to Langfuse. We use it on intent classification, extraction, and the later steps (fetch, plan, message). Each trace gets provenance metadata—git commit, branch, deploy env—so we can tie a run to a code version. We also set session IDs when we have them (e.g. from the chat UI), so we can group traces by conversation.

We also log input/output token counts and the model name on the current observation. Langfuse can then infer or display cost per call and per trace, showing which step and model is expensive during evals or production traffic.

At startup we register prompts (e.g. intent classification, information extraction) with Langfuse. That gives a single place to see which prompt versions are in use and to compare runs when we change a prompt.

2. Persisting runs and traces for later review

Evaluations are separate from unit tests: they measure performance over time and are built around Langfuse datasets and dataset runs. We keep evaluation data in git (e.g. JSONL files); a script syncs them into Langfuse as datasets. Each evaluation run loads a dataset, runs the task per item, applies evaluators, records scores on each trace, and links traces to dataset items. So in Langfuse we get one trace per test case, with input, output, and scores.

Evaluators return structured scores (e.g. intent match, correct tools selected, resolution helpfulness). We record them on the trace and also compute run-level aggregates—accuracy, pass rate, mean latency—and attach them to the dataset run. That way we can compare runs side by side: “this run after the prompt change” vs “last week’s run.”

We don’t throw away runs. We can open any past run, drill into a failed case, and see the full trace: which step failed, what the model saw and returned, and what the scorer said.

3. Turning failures into targeted fixes

Persisted traces become useful when they feed back into the codebase. A single Cursor command (/langfuse-triage-failures) runs a triage loop with three specialized subagents.

flowchart TD
    RUN(["Eval run"])
    D["🤖 eval-discovery agent<br/>list_failing_traces.py"]
    A["🤖 eval-trace-analyzer agent<br/>classify root cause, fix or no-op,<br/>comment on the Langfuse trace"]
    R["🤖 eval-rerunner agent<br/>rerun_failed_evals.py"]
    OUT(["Fixes confirmed"])

    RUN --> D
    D -->|"FailingTrace[]"| A
    A -->|"next trace, one at a time"| A
    A -->|"AnalyzerResult[]"| R
    R -->|"RerunResult[]"| OUT

    classDef agent fill:#f3e8ff,stroke:#7c3aed,stroke-width:2px,color:#1f2937
    classDef terminal fill:#f8fafc,stroke:#64748b,stroke-width:1.5px,color:#1f2937
    class D,A,R agent
    class RUN,OUT terminal

The analyzer agents run one trace at a time because each may modify source files. The orchestrator hands off the next trace only after the previous one has returned its AnalyzerResult, avoiding conflicts between parallel edits.

Each handoff carries a typed payload—FailingTrace from discovery, AnalyzerResult from the analyzer, RerunResult from the rerunner—defined in a single contracts file that all agents reference. This keeps the pipeline consistent when agents are swapped or updated.

Root-cause categories

The analyzer classifies each failure into one of seven categories, each with its own remediation guide:

CategoryWhat it meansWhere to fix
eval_data_wrongGround truth or test fixture is incorrectUpdate eval JSONL
mock_server_mismatchMock doesn’t match the real API contractFix mock server response
missing_mock_get_endpointGET endpoint missing from mockAdd endpoint to mock
missing_mock_post_endpointPOST endpoint missing from mockAdd endpoint to mock
prompts_behaviorModel does wrong thing due to prompt wordingEdit prompt template
system_behaviorCode-level bug independent of promptsFix application code
something_elseLikely flakiness or LLM non-determinismRe-run; document if recurring

For something_else the agent posts a comment on the Langfuse trace and returns implemented_fix: "no_code_change"—no code is touched. Everything is still logged so we can spot recurring patterns.

The analyzer is prohibited from editing raw customer-interaction fixtures: they are real inputs, and editing one would make a test pass without making the system any better. If a raw case fails because the message itself is messy, the fix goes into the prompt or comparison logic, not the fixture file. That boundary is written into the agent’s system prompt rather than left to convention.

4. Fix eval data before prompts

Some of our eval data is a set of patterns in JSONL files: ground-truth patterns for outputs that count as correct, and false-positive patterns for outputs known to be wrong. A model output that matches neither is an unknown positive.

Each failure still requires a choice: fix the evaluation data (ground truth and false-positive patterns) or the model behavior (prompts and system behavior). Mixing them wastes time and can make the metrics less trustworthy.

Fix eval data first. If the ground truth is wrong or the pattern matching is too strict, any metric you compute is misleading. You might “improve” precision by tightening a prompt when the real fix is correcting a pattern that counts correct outputs as wrong. One sign is a high unknown-positive count: many model outputs match no pattern at all.

Fix prompts after the eval data reflects what the model should do. Low precision or recall then points to the prompt: the model consistently misses a class of real issues (false negatives) or flags things it shouldn’t (false positives), while the patterns are already correct.

In practice this means running two distinct workflows:

  • Improve eval data: Fetch traces with unknown positives, classify each one as true positive or false positive, update the JSONL pattern files, re-run to confirm counts moved correctly. Do not touch prompts here.
  • Improve metrics: With eval data stable, analyze failure patterns (grouped false negatives and false positives), edit the relevant prompt template, re-run and compare precision/recall/F1 before and after.

Keeping these separate also prevents a common trap: adding ground-truth patterns to “fix” metrics instead of actually improving the model. If you do that, your eval data drifts away from reality and future metrics become meaningless.

What closes the loop

Together, these make a failed evaluation traceable from production context to root cause, code change, and rerun. More importantly, they keep a bad label from being “fixed” with a prompt change, or bad model behavior hidden by changing the label.

AI Chat

Messages you send are processed by the Google Gemini API to generate responses. Do not share sensitive personal data. See the privacy policy for details.