Release MeasureOperated by Reality Contact, LLC

AI evaluation and regression

Identify which cases changed before approving an AI release.

Release Measure compares deterministic checks, human labels, and model judges on ten supplied cases. The paid system turns the accepted policy into a versioned regression gate and release report.

Example result

Calibration sample: ten release cases

meter
  1. Deterministic8 pass / 2 schema holdsExact checks resolve before any model judge runs.
  2. Human review7 pass / 1 hold / 2 disputedIndependent labels preserve disagreement for adjudication.
  3. Judge policy1 false pass / 2 false holdsErrors remain visible by direction and case.
  4. ThresholdHold candidate pending two repairsThe buyer approves the threshold and release action.
Example values only. Customer scores and thresholds use the customer's approved cases and evaluator policy.

A score is weak evidence when evaluator prompts, models, or criteria change between runs.

A prompt, model, retrieval, or tool change can alter product behavior while a changed judge prompt alters the score. Without versioned cases, independent labels, and raw per-case results, the team cannot tell which system moved.

Release Measure begins with ten cases because a small calibration table can expose false passes, false holds, ambiguous criteria, and disagreements that an aggregate score conceals.

The free report shows the evaluator's errors beside the product's.

After private intake, the ten cases run through existing deterministic checks and judges, then compare with the supplied human labels. The report preserves each case, disagreement, false pass, false hold, and the assumptions behind a proposed threshold.

The paid continuation versions the gold set and policy, implements the runner and CI gate, and produces a reproducibility manifest. The customer approves the scoring policy and makes every release decision.

What comes back from the record

A report compares deterministic checks, human labels, and judge scores across ten supplied cases, showing disagreement, false passes, false holds, and a proposed release threshold.

Turnaround: The report arrives within four business days after ten readable cases, current evaluators, and release criteria are available.

The implementation turns approved cases into a release decision record.

  1. Calibrate ten cases

    Supplied cases, labels, checks, and judge results are compared to find disagreement and ambiguous release criteria.

  2. Version the policy

    The accepted cases, deterministic assertions, judge prompt, human-review sample, thresholds, and hold rules become reviewed artifacts.

  3. Implement the gate

    A regression runner compares the accepted baseline with a candidate and preserves raw results, versions, cost, latency, and failures.

  4. Approve or hold

    You inspect the report, approve the policy and threshold, and decide whether the candidate can be released.

Why the check is free

The calibration report is free because Reality Contact, LLC is measuring whether evaluation disagreements create enough release risk for teams to request an implemented regression system.

Free ten-case calibration report

A report compares deterministic checks, human labels, and judge scores across ten supplied cases, showing disagreement, false passes, false holds, and a proposed release threshold. The report arrives within four business days after ten readable cases, current evaluators, and release criteria are available.

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

Questions before you send anything

What do I send?

Do not send private files or links through this public form. If the evaluation fits, a person will reply with a secure intake method and written deletion terms before you share traces or labeled cases.

What comes back for free?

A report compares deterministic checks, human labels, and judge scores across ten supplied cases, showing disagreement, false passes, false holds, and a proposed release threshold. The report arrives within four business days after ten readable cases, current evaluators, and release criteria are available.

Where does the service stop?

Reality Contact, LLC implements and calibrates the evaluation system, but does not decide product policy, certify model safety, label regulated outcomes, approve deployment, or replace the buyer's domain experts and risk owners. The buyer approves the gold cases, scoring policy, threshold, and hold rules, then replaces the informal release check with the accepted suite for the next AI change. This is technical evaluation implementation; it does not replace domain-expert, legal, regulatory, security, safety, or product-policy review.

Can the service create our domain labels?

The service can structure and compare supplied labels, but the customer provides or approves the domain expertise and resolves policy disagreements.

Does a passing suite establish model safety?

No. A passing result describes the versioned cases, evaluators, and run policy in the manifest; it is not a safety certification or a substitute for the customer's risk review.

Can the system use our current evaluation platform?

Yes, when the platform exposes the traces, datasets, evaluators, and runner hooks needed for the agreed reproducibility and release evidence.

Free ten-case calibration report

A report compares deterministic checks, human labels, and judge scores across ten supplied cases, showing disagreement, false passes, false holds, and a proposed release threshold. The report arrives within four business days after ten readable cases, current evaluators, and release criteria are available.

Do not send private links or files through this form. If the service fits, a person will reply with a secure intake method and written deletion terms before you share private material.

Operated by Reality Contact, LLC.

The customer approves the evaluation policy, threshold, hold rules, and release decision.

Refund conditions appear beside the regression-system price.

First-party pseudonymous attention analytics · Privacy and opt-out