← The Torre work01 / TORRE · CONTRIBUTION

Tracing the calls inside an AI operation.

Make AI operations easier to investigate.

Where this work fits

Once AI features were in use, a disappointing result could come from several places. A model call, a search result, and a dashboard count each needed different evidence before the team could decide what to fix.

MY CONTRIBUTION

I added model tracing with OpenTelemetry and Langfuse. Recording timing, failures, and generation context let engineers inspect individual calls inside an operation, including failed attempts that could be hidden by a successful fallback.

THE IDEA, ILLUSTRATED

Make the path of a request inspectable

01 / MODEL BEHAVIOR

The failure leaves a trail.

One request. Every model attempt.

request / demo-1042Generation trace
Generation · attempt 1

The failed provider attempt stays visible, including its exception.

The first attempt fails; the fallback succeeds. Keeping both in the trace reveals a failure that the final response alone would hide.

THE ENGINEERING DECISION

Capture the right context without blocking the user.

The review team needed the result the user actually saw. I worked on carrying that context into the review handoff while keeping a delivery failure from interrupting search. For metrics, I checked what each count represented and kept incomplete periods from distorting comparisons.

The skills behind the work

OpenTelemetry
Instrument operations so engineers can inspect timing and behavior across execution steps.
Langfuse
Inspect model activity and the context needed to investigate AI behavior.
Asynchronous work
Let longer-running operations progress outside an immediate user request.
THE WIDER RESULT

Engineers could inspect model failures and timing. Reviewers received search context to investigate, and product reporting used clearer definitions for the activity being measured.

READ THIS IN CONTEXTI gave the team more to investigate than a failed result. →
Related contributionsConnecting product events to their search contextI instrumented recruiting actions such as reviewing a candidate, changing search criteria, and starting a conversation. Keeping search context with those events made the activity more useful for analysis than unrelated counts of clicks.Metrics with consistent definitionsI worked on SQL reporting for candidate activity, including the populations and time periods behind the numbers. I separated measures of different actions and adjusted comparisons so incomplete periods did not make a rate look worse simply because the day was still underway.Search context for the n8n review workflowI built the application-side integration that supplied the review team’s n8n workflow with search context and the candidates actually displayed. I worked on correlating that information and handling delivery failures without interrupting search. The review team maintained the receiving workflow.