RAKSHITA DEVURKAR / SOFTWARE ENGINEER

Austin, TX · Expedia Group

I build agent systems
that work beyond the demo.

At Expedia Group I work on the agent systems behind traveler support, and on how we tell whether a change to one actually helped.
2.5M conversations a month — the interesting part is what happens when one goes wrong.

Explore my work
DISTRIBUTED SYSTEMS AGENT RELIABILITY DEVELOPER TOOLS

01 / SELECTED WORK

Three problems,
and what they cost.

Escalation that loses context, improvements nobody could measure, and a bottleneck seven hundred engineers were working around.

01

AGENT SYSTEMS

Keeping context when it matters most.

Designed reusable escalation flows for the agentic Help Center, preserving traveler and booking context across tools. Redesigned support handoff around persisted context and explicit escalation reasons.

Orchestration Failure recovery Human handoff
2.5M+ monthly conversations supported
52% less human-agent handling time
02

EVALUATION

Making “better” measurable.

Built a multi-turn evaluation framework combining deterministic evaluators with calibrated LLM judges. Validated judgments against 1,400 human-labeled samples to make tool-selection improvements measurable.

Multi-turn evaluation LLM judges
0.67 → 0.91 tool-selection accuracy
>90% agreement with human experts
03

DEVELOPER EXPERIENCE

A recurring bottleneck. A shared tool.

Engineer interviews surfaced test-booking generation as a recurring obstacle. Shipped a CLI prototype in four weeks, then productized booking APIs as versioned MCP tools for adoption across the organization.

MCP API design Internal platforms
700 engineers using the tools
4 weeks from problem to CLI prototype

CASE STUDY / ESCALATION CONTEXT

Why we persisted context instead of re-deriving it.

The problem

When the agentic Help Center handed a conversation to a human, the human started cold. The traveller had already explained the problem, the agent had already looked up the booking, and none of it survived the handoff. Travellers repeated themselves to a second person who could see less than the first.

The constraints

  • Regulatory. Booking and traveller data could not be copied into a new store without a retention story, so "log everything and let the agent console read it" was not available.
  • Existing tools. The agent's context was spread across several tools owned by different teams. No one team could change the shape of it unilaterally.
  • Live traffic. Roughly 2.5M conversations a month, so migration had to be incremental. There was no cutover.

The alternatives

  • Re-derive on handoff. Have the human console call the same tools again. Cheap to build and needed no new storage — but it doubled tool load at exactly the moment a conversation was already going badly, and it could not reconstruct why the agent had escalated.
  • Pass a transcript. Hand the human the conversation and let them read it. Simple, and what most handoffs do — but it moves the work to the person, and a transcript does not say which booking was in scope.
  • Persist a context record, keyed to the conversation. More to build, and a retention question to answer, but the escalation reason becomes explicit rather than inferred.

The choice, and the trade

We persisted a context record with an explicit escalation reason. The deciding argument was not effort: it was that the first two options make the handoff a guess. A human who can see the agent tried a refund and hit a policy limit is solving a different problem than one reading a transcript.

The cost we accepted: a new record with a retention policy, and the escalation reason becoming an interface that other teams now depend on — which means it cannot be changed casually.

Bringing other teams along

[TODO — the part only you can write. What actually moved the tool-owning teams: was it a review, a prototype that showed the console with context in it, one incident that made the cost concrete? Name the specific thing and the specific objection you had to answer.]

What I would do differently

[TODO — one honest thing. The escalation-reason taxonomy being harder to agree than the storage? Underestimating how many callers would depend on it? Whatever you actually got wrong.]

2.5M+monthly conversations
52%less human handling time

02 / OPEN SOURCE

Two tools I built
to check my own work.

Two projects exploring different parts of the agent evaluation loop.

01 / REGRESSION COVERAGE Python

evalkeep↗

Turn agent failures into reviewed regression tests. Keep the evidence, preserve human judgment, and check whether a fix actually holds.

pip install evalkeep

One failure, becoming a test
  1. A trace — a customer asks to refund their latest order. The agent refunds the oldest one.
  2. A family — 8,562 failures group into 239 families; this one covers 1,096 of them, so it is worth a test.
  3. A draft — refund_order.order_id != 'order-A', with the tool results the original agent saw attached as fixtures.
  4. A decision — nothing is exported until a person approves it. The draft says outright that it forbids the observed mistake without knowing the right answer.
02 / REPEATABLE EXPERIMENTS Python · AWS

Agent Trial Runner↗

Run the same task repeatedly in isolated environments. Grade the state an agent leaves behind, and separate wrong answers from trials that could not finish.

2,500 tasks × 12 configs × 10 repeats Designed for 300,000 trials · 5,040-trial dry run documented

03 / FIELD NOTES

Things that surprised me.

Design decisions, inconvenient results, and the details hiding behind a score.

01

EXPERIMENTS · 7 MIN READ

What I learnt when I ran
5,040 agent trials.

Prompt sensitivity, flaky success, and why the grader deserved its own investigation.

02

EVALUATION DESIGN · 6 MIN READ

A production failure should
leave behind a test.

Building evalkeep, and the gap between detecting a mistake and defining the right behavior.

04 / A LITTLE CONTEXT

I like the parts
that connect.

A user’s problem and a system’s behavior.
A promising prototype and a tool a team relies on.
A score and the evidence underneath it.

I’m Rakshita, a software engineer at Expedia Group in Austin. My work spans distributed backends, full-stack products, and agentic AI. I’m drawn to problems that need both careful implementation and a view of the whole system.

Beyond shipping, I mentor engineers and lead cross-team architecture reviews around evaluation, rollout, observability, and recovery. Before Expedia, I built shared frontend infrastructure at MeridianLink and messaging systems at MetricStream.

2022 — NOW Expedia Group SDE III
2018 — 2022 MeridianLink Software Engineer
2017 Onjax Technologies Web Development Intern
2015 — 2016 MetricStream Member of Technical Staff

M.S. Computer Science · Binghamton University
B.E. Information Science · RNS Institute of Technology

Read my resume ↗

KEEP THE CONVERSATION GOING

What are you
working on?

rakshitadevurkar@gmail.com

Always happy to compare notes on agents, evaluation, and building useful software.