rudvanth-eval
PrototypeThe evaluation harness we use on every engagement. Declarative suite definitions, pluggable graders, CI-friendly output, and regression tracking across model and prompt changes.
Licence · Apache-2.0
, Open source
01Release plan
These are commitments with dates attached internally. Status here mirrors reality, an entry moves to a public repository only when there is something worth cloning.
The evaluation harness we use on every engagement. Declarative suite definitions, pluggable graders, CI-friendly output, and regression tracking across model and prompt changes.
Licence · Apache-2.0
A benchmark for citation faithfulness: does a generated statement actually follow from the passage it cites? Includes an annotation protocol so others can extend it to their own domain.
Licence · CC BY-SA 4.0
Test sets for calibrated refusal, measuring whether a system stays quiet when the corpus does not contain the answer, which is the behaviour enterprise buyers care most about and public benchmarks measure least.
Licence · Apache-2.0
Openly licensed domain evaluation sets for Indian-language and code-switched enterprise text, starting with healthcare and financial services terminology.
Licence · CC BY 4.0
02Policy
Anyone can say their retrieval is accurate. A harness someone can run against their own corpus is a different kind of statement, because it can be checked. That is the only kind of claim worth making about engineering rigour.
Code under Apache-2.0, data and benchmarks under Creative Commons. We are not trying to build a moat out of a licence, the moat is knowing how to use the thing.
A repository with no commits for a year and eleven open issues is worse for our credibility than no repository. Anything we stop maintaining gets archived with a note saying so.
Nothing derived from a client corpus is published in any form, including as synthetic data generated from it. Benchmarks are built from public or purpose-created material.
Issues and pull requests are welcome on every public repository. The contributions we value most are new hard cases for the evaluation suites, a document, a query and an expected behaviour that current systems get wrong.
We also review security reports through a coordinated disclosure process.
Research enquiries
[email protected]Security reports
[email protected]