Tracing the calls inside an AI operation.
Make AI operations easier to investigate.
Where this work fits
Once AI features were in use, a disappointing result could come from several places. A model call, a search result, and a dashboard count each needed different evidence before the team could decide what to fix.
I added model tracing with OpenTelemetry and Langfuse. Recording timing, failures, and generation context let engineers inspect individual calls inside an operation, including failed attempts that could be hidden by a successful fallback.
Make the path of a request inspectable
The failure leaves a trail.
One request. Every model attempt.
request / demo-1042Generation traceThe failed provider attempt stays visible, including its exception.
The first attempt fails; the fallback succeeds. Keeping both in the trace reveals a failure that the final response alone would hide.
Capture the right context without blocking the user.
The review team needed the result the user actually saw. I worked on carrying that context into the review handoff while keeping a delivery failure from interrupting search. For metrics, I checked what each count represented and kept incomplete periods from distorting comparisons.
The skills behind the work
- OpenTelemetry
- Instrument operations so engineers can inspect timing and behavior across execution steps.
- Langfuse
- Inspect model activity and the context needed to investigate AI behavior.
- Asynchronous work
- Let longer-running operations progress outside an immediate user request.
Engineers could inspect model failures and timing. Reviewers received search context to investigate, and product reporting used clearer definitions for the activity being measured.