The Model Never Produces a Number: Vinoo Ganesh on Building Ground Truth into AI for Finance

Q1. At Kepler, your team focuses on “separating what AI does well from what code does well” to ensure absolute ground truth. In the context of financial services—where a misplaced decimal point can break a model—how do you structurally design an infrastructure where LLMs handle unstructured synthesis, but deterministic code handles computation and verification? Where do you draw the architectural boundary line?

The rule we hold to is that the model never produces a number, and most of our architecture is built in service of that.

For example, consider a question like “what was gross margin last fiscal year, excluding the services segment.” An analyst asking it could mean several different things, depending on which fiscal calendar the company reports on, which entity in the corporate structure they care about, and whether excluding a segment means removing both its revenue and its cost or its revenue alone. Resolving that kind of ambiguity is language work, which is where LLMs are strong, so we give the model the job of turning the question into a structured plan that names the entity, the metric, the period, any adjustments, and the sources the system is permitted to use.

Once the plan exists, the model’s role in the computation is over. The plan compiles into deterministic operations that retrieve typed facts from our store and apply the arithmetic with explicit handling of units, scale and currency, and every intermediate value carries its lineage forward. When the model writes the final response, it writes around placeholders that the engine fills with computed figures. A verification pass then confirms that every number in the text traces back to a value the engine produced, and an answer that fails that check is never shown to the user.

The boundary therefore falls between tasks that have many acceptable answers and tasks that have only one. Interpreting intent, choosing the applicable definition and phrasing the result belong to the model, while retrieval, arithmetic and period alignment belong to code. Where the model cannot resolve intent with confidence, it asks the analyst a clarifying question rather than guessing.

The design has real costs. Each question pays for a planning step before anything executes, which can make the system slower than a single model call, and some questions don’t reduce to a plan at all. A question about how management described demand on an earnings call has nothing to compute, so the model answers from the relevant passages, quotes them, and the response is labeled as qualitative (though it still has full citations) so that no one mistakes it for a verified figure.

What surprised us was how much analysts value the plan itself. They read it, challenge it and correct it before any number is computed, and it has become the part of the product they trust most. It is also what allows us to remain model-agnostic, because a model that never held the truth can be replaced without changing a single figure.

Q2. One of Kepler’s core promises is perfect provenance, making every figure traceable back to its source filing, page, and line item. Given your deep background managing Spark/YARN compute platforms and indexing backends at Palantir, what does the underlying data and storage layer look like to support this depth of real-time, bi-directional traceability? How does it differ from traditional vector databases or RAG architectures?

In a typical RAG setup, documents are chunked and embedded, and the model is handed whichever chunks look most similar to the question. So, the best it can offer afterwards, is a passage where the answer probably lives. That breaks down in finance, where the 2023 and 2024 revenue tables look nearly identical to an embedding model, and where a citation to roughly the right page isn’t the same as knowing which number was used and why.

We make provenance a property of the data instead of something attached to the answer at the end, which means the work happens at ingestion rather than at query time. Our extraction models read every page of every document, scanned PDFs included, and store each fact with the header, period and unit that give it meaning, along with its location in the source (among others). Companies describe the same line item in different ways and change their wording over time, so specialized models map each company’s labels to a standard taxonomy, which is what makes a fact from one filing comparable to a fact from another.

What we realized building this is that deep preprocessing matters far more than retrieval cleverness. This means the concerns are ones you’d recognize from any serious analytics engine. Partitioning is the obvious one, since facts organized by entity, period and statement type mean a question about one company’s last eight quarters touches a narrow slice of the corpus rather than scanning across millions of documents. We compute and store statistics alongside the data, so the engine knows what exists before it goes looking and can prune most of the search space up front. The third is schema completeness, because a fact missing its scale or its currency is worse than a fact that isn’t there at all, and a partial schema will silently produce a number that’s off by a factor of a thousand. Every extraction is validated against the schema before it lands, and anything incomplete is quarantined rather than written. Access controls apply at retrieval too, so a user only ever pulls from sources they’re entitled to see.

Traceability then runs in both directions. From any number, an analyst can click through to the line item highlighted in the original filing, including inside the Excel templates their desk already uses. In the other direction, every computed figure keeps its derivation chain, recording which facts went in and which formula from the ontology was applied, and because the pipeline is idempotent, any derivation can be replayed with the same result.

Don’t get me wrong, embeddings still play a role in locating the right document or concept, but they never return a value. That separation, plus the preprocessing underneath it, is what lets us cover more than 26M+ SEC filings across 14,000+ companies and 27 markets without losing the link between a figure and its source.Q3. 

Q3. As the former CTO of Veraset, you managed geospatial data pipelines processing over 2+ TB of daily data. Big Data engineering traditionally relies on structured pipelines and predictable data flows. When moving from massive-scale deterministic data ingestion to the probabilistic world of modern AI, what are the hardest infrastructure engineering lessons you’ve had to port over—or completely unlearn?

At their core, and in a very simplified way, structured pipelines do three things. They ingest data with a known shape, transform it through stages you can reason about, and write it somewhere with a contract attached (this is mostly a variant of ETL). Everything I built for years assumed that shape held, and when it didn’t, I found out fast. A job OOMs, a stage throws an exception, a partition lands empty, and PagerDuty wakes me up. Those failures are the good ones. A pipeline that breaks hard enough to stop processing keeps bad data from traveling downstream. What actually hurt me was the other kind. An upstream provider changes an SDK or a device panel shifts underneath me. My jobs keep running, my health checks stay green, and bad data moves through every stage untouched and into the customer’s hands. Three weeks later someone asks why a number looks off (this happened more than I care to admit).

A model returns a fluent, confident, wrong answer and throws no exception, which puts it in that second category. The complication is that this isn’t a bug I can fix. It’s a feature I depend on, since the entropy producing the wrong answer is the same entropy that makes the model useful for interpretation. So we test the data rather than the infrastructure around it, and we treat model output as an untrusted upstream feed. Every prompt change, model version, and context change runs against thousands of cases with known answers before reaching production. We check the plan the model produced alongside the final computed result.

A footnote telling me a segment was reclassified in the prior year carries an entity, a period, an affected line item, and a direction. For most of my career I would have called that unstructured and sent it to an unstructured store, because structure meant a schema and that text doesn’t fit one. But much of what moves a financial answer lives in exactly that kind of prose. Inferring what the source never explicitly typed is the hard part. We use models for this, which makes the schema a hypothesis rather than a contract, and makes schema evolution the normal state of the system rather than an anomaly. A firm has its own definition of adjusted EBITDA, which evolves when they acquire something. An instrument appears that my capital structure logic has never seen. None of that should force a migration. Our ontology carries the load as an extensible layer, with concepts, formulas, and relationships we version and extend per customer while the facts underneath stay untouched.

When types are enforced, the wrong shape can and does blow up the job. When types are inferred, the wrong shape becomes a fact with a plausible unit and a plausible period, sitting in the store waiting for someone to compute with it. We catch that at ingestion, quarantining anything that fails its schema instead of writing it. The damage happens at write time, and nothing downstream will ever report it. Preprocessing therefore earns more of my attention than retrieval does. This mirrors how I used to think about Spark, where I cared about how a dataset was partitioned, what statistics sat alongside it, and whether the schema was complete enough to prune work before a job ever ran. At Kepler, we partition facts by entity, period, and statement type, so a question about eight quarters at one company touches a narrow slice rather than scanning the entire corpus. We store statistics next to the data so the engine knows what exists before it goes looking. And we refuse to write a fact missing its scale or its currency, because a partial schema produces a number off by a factor of a thousand that reads as perfectly reasonable.

Our extraction follows the same logic. Handing a frontier model a PDF works, but it costs a fortune and gives me nothing I can audit. Meanwhile, a 10-K, an earnings deck, a credit agreement, and a scanned fund statement each carry their own conventions about where meaning sits, which header governs which column, how periods get labeled, where units are declared. So there is structure. It’s just document specific. We run small specialized models per document type, and bring the frontier model in where the input is a person’s question rather than a document we already understand.

Q4. Your team is stacked with alumni from Palantir, Citadel, Meta, and Stanford, and backed by foundational data/AI pioneers (OpenAI, dbt, MotherDuck). Having built high-stakes compute platforms for defense and quantitative finance, what architectural methodologies from your past are you embedding into Kepler to ensure the platform meets “mission-critical” availability and correctness standards?

Sensitive data doesn’t leave a customer’s environment, so Kepler deploys into siloed, customer-hosted environments rather than asking a firm to ship its documents to us. Access controls apply at every step of the pipeline, which means a user retrieves only from sources they’re entitled to see, and the same restriction holds whether the request comes from the interface, an Excel template, or a connector. Audit logging covers the full path of every request. We built all of this in from the start rather than adding it during a procurement cycle, and we’re SOC 2 Type II certified with ISO 27001 underway.

Correctness comes from determinism, which sounds abstract until you watch what happens without it. The same question has to produce the same answer on Tuesday that it produced on Monday, so our workflows are idempotent by design. An enterprise value build across a complex capital structure, with preferred shares, convertibles, and minority interests, runs as the same sequence of deterministic operations every time. Reconciling segment revenue when a company changes how it reports works the same way. If an answer can’t be replayed, it can’t be defended to a client or an auditor, and that constraint shapes more of the architecture than anything else.

Evaluation functions as a gate rather than a report. Every stage of the pipeline is tested on its own and end to end against known correct answers, covering the plan the model produced as well as the final computed figure. Nothing reaches production without passing. When a test does fail, we can tell whether the problem sits in the model’s reasoning, in the context we supplied, or in the downstream execution, which makes the failure fixable rather than mysterious. A silent regression is how you lose a client permanently in this industry, so the cost of running thousands of cases on every change is easy to justify.

Availability follows from the modularity. Different models handle different stages depending on how much reasoning each one demands, and the deterministic layer doesn’t depend on any single model or provider. A provider outage or a model change affects how a question gets interpreted, never what the numbers are, and we can reroute or upgrade one stage without touching the rest. When the system can’t answer with confidence, it says so. That’s the right behavior here, even though it makes for a less exciting demo.

Q5. You have written extensively about Forward Deployed Engineering (FDE) and why many companies misclassify it as a mere consulting or pre-sales function. As a fellow in a16z’s FDE program, how are you leveraging the FDE model at Kepler to deploy deterministic AI systems into deeply entrenched corporate architectures? How does having engineers on the front lines change the way your core platform infrastructure is built?

At Kepler the forward deployed function reports into product rather than sales. We structured it that way before we had the customer volume to justify it, because the reporting line determines what the engineers optimize for. If you point the function at sales, the incentive becomes closing the account in front of you, which is incredibly important, but not what I see in the remit of an FDE. To me, FDEs represent an extension of the product function. They embed with customers, understand disparate problems, and generalize solutions to those problems into the core product. That is what gives the company product leverage. Pointing them at product means every deployment is expected to produce something the next deployment can start from, which is also what separates the role from consulting, where you learn one company’s model, ship something shaped exactly to it, and lose all of it when the engagement closes.

The way that plays out in practice is that our engineers go in to collect nouns and verbs, meaning the objects a business treats as real and the system of operations they follow when making decisions. Every firm we work with is running some version of the same ontology underneath, and every one of them describes it differently. A position means one thing on a credit desk and something adjacent on an equities desk inside the same bank, and two funds will use identical language for a return calculation while disagreeing about what belongs in the denominator. Most of those differences exist because the differences are what makes a firm unique. In other words, they represent their alpha. You cannot infer that from a discovery call, and you cannot really ask for it either, since the people holding that knowledge do not know they have it. A schema will tell you what is stored, but it will never tell you what is meant, and the distance between those two is where a system that sounds right produces a number that is wrong.

Deploying Kepler into an entrenched architecture is done in service of resolving that distance, and the determinism is what makes it tractable. Every institution we sell to shares one non-negotiable trait. Numbers have to be right, and somebody has to be able to show why. Designing against that constraint forces the operating model into the open rather than letting it stay implicit. A system that improvises around a definition it does not understand, or that is non-deterministic about it, will never tell you it misunderstood. Ours does not improvise, so when we get a firm’s definition wrong, that surfaces as a failure instead of as an answer that merely looks reasonable, and the engineer who got it wrong hears about it from the system that week rather than from a client six weeks later.

Those corrections are how we learn in the platform, though the signal is actually pretty narrow. We are not trying to learn which features a given fund would like to have, since three firms asking for the same feature is easy to notice and worth relatively little. What matters is where the platform is too narrow to generalize what we keep running into, and that usually arrives quietly – as an engineer working around the same limitation for the third time. Most of what comes back is built into the core ontology, since it is the layer built to absorb one firm’s definitions without forcing a migration on everyone else, and the rest shows up as workflows that get reused across deployments. The judgment that takes longest to build is knowing which fixes belong in the platform and which ones you throw away on purpose, because every shortcut you ship is something you end up owning. I learned that expensively when a script I hacked together one afternoon was still running a year later across a customer of nearly a hundred thousand people, with my name fused to it (vinoo.groovy).

Our goal is product leverage, but some people call this compounding. Every firm we map makes the next deployment cheaper to attempt, and cheap attempts are how a small company learns anything quickly. It is also where I would put the moat in this era, because it is not the model, which cheapens by the month and which you are renting from somebody else regardless, and it is not the map of any one customer, since extraction is nearly free and anyone can draft how a firm operates in an afternoon. The holy grail is the system of operations, and access to that is earned and experienced, not given. 

…………………………………………………..

Vinoo Ganesh is CEO of Kepler, the industry-agnostic, ground truth platform for AI, starting with finance. At Kepler, language models interpret financial questions and plan the analysis, while deterministic software retrieves source data, runs the calculations, and records how every result was produced. Kepler is backed by founders of OpenAI, Facebook, MotherDuck, DBT, Square, and others.

He started his career at Palantir, where he spent seven years working on the search and indexing backend, managing the Spark/YARN compute platform, and leading customer engagements across the commercial, healthcare, and defense sectors. He also led Project Frontline, Palantir’s FDE rotation program. After Palantir he served as CTO of Veraset, a geospatial data startup that grew to $15M ARR before it was acquired. Before founding Kepler, he was Head of Business Engineering at Citadel, overseeing data pipelines and investment platforms across the hedge fund.

You may also like...