Skip to content

, Applied AI research & products

AI for domains where approximately right is not good enough.

We build intelligent software, AI infrastructure and enterprise systems, while advancing long-term research into the technologies that make those systems trustworthy. Evaluation-first, built to be verified.

01What we do

Most AI projects do not fail because the model is not clever enough. They fail because nobody could say, with evidence, whether the system was working.

That is the gap Rudvanth works in. We build AI systems for organisations that have to answer to a regulator, an auditor, a clinician or a court, where a confident wrong answer is worse than no answer at all.

In practice that means starting from the measurement problem. Before we write the feature we write the evaluation set: the hard cases, the failure modes, the thresholds the system has to clear before anyone is allowed to depend on it. The model is the easy, replaceable part. The harness is the asset.

The company is organised around pillars that feed each other. Research produces methods. Products turn methods into software we own. Engineering drags the research back to reality every time it drifts.

03Why we exist

The mission is measurement.

To make AI systems that can be verified, in domains where being approximately right is not good enough. Everything else the company does is a means to that.

Mission

Build the measurement layer enterprise AI is missing. Turn that knowledge into products we own. Publish the methods, so every claim we make can be checked.

Vision

In ten years, putting an untested AI system into a hospital, a bank or a courtroom feels as reckless as shipping untested code feels today. And when someone asks how testing AI became the norm, the answer points to instruments this company built.

Every claim we make can be checked. Every system we ship carries the evidence of its own performance.

04Operating principles

Four commitments we are willing to be held to.

These are not values-page decoration. Each one costs us something specific, which is the only reason it is worth writing down.

01

Evaluation before capability

A system that cannot be measured cannot be improved, defended, or safely handed to a customer. We build the harness before we build the feature, and we keep the harness in the repository next to the code it grades.

02

Own the hard part

Wrapping a model API is not a business. We take the work that is genuinely difficult, the retrieval layer, the evaluation set, the domain constraints, and we make that the part we own.

03

Publish the method

Methods go out in the open even when the product does not. Open evaluation harnesses and honest write-ups cost us nothing we needed to keep, and published methods are how technical trust is actually earned.

04

Focused teams, long horizons

We would rather do four things properly over three years than twenty things badly in one. Every commitment on this site is one we intend to still be honouring in 2030.

05Where we work

Industries where the cost of being wrong is legible.

We deliberately avoid domains where nobody can tell whether the output was any good. If the mistake is invisible, the evaluation problem is unsolvable, and the work is not interesting.

01

Healthcare & life sciences

Clinical documentation, coding and prior-authorisation workflows where an error has a named patient attached to it.

02

Financial services

Underwriting, KYC, surveillance and reporting, domains with an auditor at the end of every decision path.

03

Manufacturing & mobility

Inspection, maintenance and field intelligence, usually on constrained hardware and intermittent connectivity.

04

Energy, climate & ESG

Disclosure, measurement and assurance pipelines where the output is a number somebody has to certify.

05

Public sector & infrastructure

Citizen-facing services in multiple Indian languages, built to be inspected rather than trusted on faith.

06

Legal & compliance

Contract and obligation analysis where the value is in the citation, not the summary.

06Open source

We release the tools we grade ourselves with.

Inspectable artefacts settle questions that assertions cannot. An evaluation harness someone else can run against their own data says more about how we work than any case study we could write.

$ rudvanth-eval run --suite clinical-coding

  suite     clinical-coding        n=1,284
  model     under-test             temp=0.0

  exact-match .................... 0.913
  citation-grounded .............. 0.968
  refusal-when-unsupported ....... 0.994
  p95 latency .................... 1.4s

  ✓ 4/4 gates passed
  → report written to ./evals/2026-08-07.json

Illustrative output. Our evaluation tooling and its results are published as they are released.

Start a conversation

Tell us what has to be right.

If you have an AI system that has to survive an audit, a clinician, or a regulator, we would like to hear about it. We reply to every serious enquiry within two working days.