Release MeasureOperated by Reality Contact, LLC

Specific answer

What an AI evaluation reproducibility manifest should record

A field-by-field guide to dataset, prompt, model, tool, retrieval, grader, environment, cost, latency, and approval versions.

A reproducibility manifest lists every version and external dependency needed to rerun an evaluation and account for differences in later results.

Record the evaluated system as a complete version

The manifest should identify the application commit, prompt set, model and dated alias when available, parameters, tool contracts, retrieval index or corpus snapshot, policy files, feature flags, and runtime environment. A model name alone is insufficient when a retrieval update or changed tool response can alter the result. External services should include the observable version or capture date the run relied on.

The case set needs its own version, split, row count, content hash, provenance policy, and exclusions. Record any transformation applied before evaluation, such as redaction, truncation, normalization, or synthetic augmentation. Those steps can change what the evaluator sees and must remain reproducible without exposing private source material in the public report.

Version every evaluator and run condition

For deterministic checks, record the code commit and configuration. For model-based judges, record the prompt, model, parameters, reference context, output schema, retry policy, and calibration-set result. Human review needs the rubric version, reviewer roles, sampling method, disagreement rule, and adjudication authority. Store raw per-case outcomes rather than only the aggregate table.

The run record should include start time, concurrency, repeated-trial policy, provider errors, retries, latency, token or platform cost, and missing cases. LangSmith and Braintrust both charge in part for traces, data, or scores, which makes cost and coverage part of the operating record rather than an afterthought.

Make the final report traceable

Each chart or release claim should resolve to the manifest, case set, raw results, and code that produced it. The approval entry names the threshold policy, blocking findings, unresolved uncertainty, decision owner, and date. The manifest must let another engineer rerun the accepted suite and explain any difference from the recorded result. Archive the runner command and environment lockfile with the result bundle; a version list without an executable entry point still leaves the next engineer to reconstruct how the recorded evaluation was produced.

Release Measure assembles and verifies the manifest through Reality Contact, LLC. The customer controls private artifacts and approves retention, evaluator policy, and the final release decision.

Where the service stops

Reality Contact, LLC implements and calibrates the evaluation system, but does not decide product policy, certify model safety, label regulated outcomes, approve deployment, or replace the buyer's domain experts and risk owners. The buyer approves the gold cases, scoring policy, threshold, and hold rules, then replaces the informal release check with the accepted suite for the next AI change. This is technical evaluation implementation; it does not replace domain-expert, legal, regulatory, security, safety, or product-policy review. We do not promise coverage of every behavior, a defect-free model, evaluator agreement, successful deployment, or approval of any release.

Sources: LangSmith plans for traces, datasets, and evaluation; Braintrust plans for traces, datasets, experiments, and scores.

Free ten-case calibration report

A report compares deterministic checks, human labels, and judge scores across ten supplied cases, showing disagreement, false passes, false holds, and a proposed release threshold. The report arrives within four business days after ten readable cases, current evaluators, and release criteria are available.

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

Questions about this answer

AI evaluation reproducibility manifest checklist?

A reproducibility manifest lists every version and external dependency needed to rerun an evaluation and account for differences in later results.

What should I send for the free check?

Do not send private files or links through this public form. If the evaluation fits, a person will reply with a secure intake method and written deletion terms before you share traces or labeled cases.

What does Reality Contact, LLC do?

Reality Contact, LLC implements and calibrates the evaluation system, but does not decide product policy, certify model safety, label regulated outcomes, approve deployment, or replace the buyer's domain experts and risk owners. The buyer approves the gold cases, scoring policy, threshold, and hold rules, then replaces the informal release check with the accepted suite for the next AI change.

Operated by Reality Contact, LLC.

The customer approves the evaluation policy, threshold, hold rules, and release decision.

First-party pseudonymous attention analytics · Privacy and opt-out