Fina helps people with finance questions and tasks. Some need a quick answer; others involve reading documents, running calculations, checking assumptions, and building a workbook. A single request can take many model calls, with each step depending on what came before.
Finance teams judge that work the way they judge an analyst's: the numbers have to tie out, and the deliverable has to arrive. That makes the connection between Fina and the model part of the product. If one call in the chain loses context, stalls, or fails without a backup, the whole job suffers, and not always visibly.
LiteLLM is our AI gateway. It routes Fina's requests to different model providers and handles their differences in request formats, authentication, and responses. It also manages fallbacks when a provider is unavailable.
It was also where we found four bugs that only surfaced in real agent runs: reasoning state dropped between steps, an authentication setting lost when calling Claude through Databricks, a tool streaming option that never reached the model, and fallbacks that gave up while a healthy model was still available.
So we forked LiteLLM.
Maintaining a small fork let us fix the gateway behavior Fina depended on, while keeping LiteLLM's broad provider support. We could patch an issue, test it through a complete finance workflow, and ship the fix on our schedule. We also contribute fixes upstream so we have less code to maintain over time.
Why put a gateway in the middle?

Each model provider has its own protocol for accepting requests, returning answers, and reporting errors. If our application handled every protocol difference itself, adding or changing a provider would spread that work across the codebase.
A gateway puts much of that translation in one place. Our agent asks for a model response; LiteLLM handles the provider's particular format.
That is useful until something important gets lost in translation.
For a simple chatbot, a test might be: send a message and check that an answer comes back. An agent needs more. It asks a model to call a tool, runs that tool, sends the result back, and continues. Every part of that exchange has to survive the trip through the gateway.
That longer loop is where our problems showed up.
Carrying context into the next step
Picture an analyst halfway through a variance analysis who asks you to pull one number. When you hand it back, you expect them to pick up with the same plan, not start reasoning from scratch.
Reasoning models need continuity too. Alongside the visible answer, some return encrypted state that helps them continue after a tool call. Our application does not need to interpret that state. It needs to preserve it and send it back correctly.

During internal testing, we found a case where LiteLLM carried that information internally but dropped it when sending a streamed response to the client. The tool call could still arrive, which made the problem easy to miss. Getting an answer did not necessarily mean the next step had all its context.
In a finance workflow, that missing context can affect the steps that follow. The tool call can still arrive without an obvious error, making the gap easy to miss. Our domain experts noticed better multi-step answers once the reasoning state survived the trip.
The fix was small: make sure the code transforming the streamed response preserves the reasoning state along with the visible text.
The original fix had already been proposed by a community contributor. We adopted it into our fork and added more tests to catch this problem in our upstream submission.
Calling Claude through different providers
We call Claude through more than one provider, including Databricks. More providers give us more capacity and more places to fall back to. That adds another translation boundary: we are calling an Anthropic model through another provider's infrastructure.
To preserve the context needed across tool calls, we used Databricks' connection that speaks Anthropic's own Messages API protocol. Our application could keep its existing interface while the LiteLLM gateway used the provider connection suited to that conversation.
Sounds good in theory, but in practice, authentication exposed another gap in upstream LiteLLM.
LiteLLM already had an option for the kind of authentication this endpoint expected. But on the path we used, that option was not reaching the code that prepared the request. The right setting was present in our configuration and absent where it mattered.
Our fix forwards the setting correctly and keeps it out of the request body sent to the model.
It was a useful reminder that provider integration is only the beginning for an LLM gateway. The routes, authentication, and inference all need to work together.
When tool calls go quiet
Some finance requests end with a short explanation. Others produce a workbook or a detailed report. When Fina writes a large artifact through a tool call, it benefits from receiving the tool inputs as they are generated.
If the provider waits until it has assembled the whole tool input before sending anything, the application sees a long stretch of silence. To an analyst waiting on a deliverable, that silence looks the same as a hang.
We internally investigated pauses like this during large artifact writes, particularly with Anthropic models. Provider-side buffering was the biggest suspected contributor, so we thoroughly tested our pipeline to see exactly what reached the provider through the gateway.
As it turned out, the LiteLLM gateway was silently dropping Anthropic's fine-grained tool streaming option. We patched the fork to preserve it, allowing the setting to reach Claude.
Support for the streaming option has since been merged into LiteLLM.
When the backup fails too
Financial work runs on deadlines: a month-end close, a filing date, a board pack. Model providers frequently go down, hit capacity limits, or time out, and an agent working to a deadline cannot simply wait for them to recover. When a task needs repeated model calls, a single unavailable endpoint can interrupt the whole job. Multiple fallback models, ideally across different providers, give the agent other ways to continue. Those backups need to support the tools and context the task requires.

Streaming makes recovery a different problem. If a request fails immediately, the gateway can try the next model. But a stream can start successfully and break while the application is reading it, even after part of the answer or a tool call has arrived. Recovery has to account for that unfinished output as well as choose another model.
The bug we found was in keeping those backup choices available. Suppose model A has two backups in the fallback chain, B and C. If model A fails, LiteLLM returns a stream from B. If B then fails while that stream is being read, a working model C could be left untried, and the user would see the task fail even though a healthy model was available.
We patched the router to retain the untried alternatives for the lifetime of the stream. When a backup fails, it can return to the remaining options instead of giving up early.
Iterating fast in isolation
Our namespaced staging environments helped us move quickly here. They gave us a separate place to run Fina with our LiteLLM fork, reproduce failures, and trace them back to the gateway without touching production.
From there, we could apply a patch, deploy it into the same environment, and run the failing flow again. Keeping that loop short made it easier to identify root causes and check fixes in the full agent workflow.
Keeping the fork small
Forking LiteLLM got us past the immediate blockers and let us ship fixes on our own schedule. Keeping it up to date takes work, though. Each upgrade means checking our patches against the latest upstream code.
That's why we're contributing fixes back with tests and context from the existing issues. When LiteLLM includes an equivalent fix, we can drop our patch. The less code we have to maintain in the gateway, the more time we can spend building Fina.
For Fina, the gateway has to support the whole task: carry context into the next step, deliver tool calls correctly, and keep remaining providers available when one fails. That is how we test our LiteLLM fork—through complete finance workflows, including the points where things go wrong.







