Pre-launch counsel-review draft
AI Evaluation Limitations
How to interpret automated checks, model-assisted findings, verification, health scores, and alerts without treating them as infallible truth.
Professional legal review required before publication
This draft is not effective, has not been approved by counsel, and is not legal advice. Bracketed decisions and review markers must be resolved before publication.
- Effective date
- To be set after professional legal approval
- Last revised
- To be set after professional legal approval
- Draft record
- M27 counsel-review draft 1, prepared August 27, 2026
Use results as evidence, not unquestionable truth
NoirGen Agent QA combines system evidence, deterministic assertions, customer-provided facts and rubrics, and model-assisted evaluation to help identify potential quality regressions in text-based AI agents. These techniques can improve review coverage, but they cannot prove that an agent is always correct, safe, secure, compliant, or suitable for a particular use.
- Model-assisted evaluation is probabilistic and can produce false positives, false negatives, incomplete explanations, or inconsistent results.
- Deterministic checks reduce uncertainty only for the explicit assertions and inputs they cover.
- High-severity verification is an additional check, not a guarantee that a finding is correct or complete.
- Scores and findings depend on customer-provided prompts, rubrics, facts, test coverage, agent behavior, and provider behavior.
- Models, prompts, external services, and customer agents can change over time, so prior results may not predict future results.
- Results support quality assurance; they are not legal, regulatory, medical, financial, security, or other professional advice.
- Customers must review important findings and remain responsible for decisions, deployments, and communications based on the results.
These limitations form part of the proposed Terms of Service and should be read together with the Acceptable Use Policy.
1. Different layers answer different questions
The evaluation pipeline separates several kinds of evidence:
- Execution validity asks whether the configured endpoint returned a bounded, usable response. A successful request is not by itself a quality pass.
- Deterministic assertions apply explicit checks such as contains, excludes, equals, regular expression, length, or valid-URL rules.
- Knowledge-grounded and rubric evaluation uses the supplied facts, expected behavior, concepts, and rubrics to assess eligible responses.
- High-severity verification performs an additional model-assisted check for critical, ambiguous, or configured high-severity concerns.
- Health scoring and regression comparison aggregate stored evidence and compare eligible runs under versioned rules.
Each layer has a bounded purpose. One layer cannot repair missing test coverage, incorrect customer facts, an unavailable endpoint, or an error produced by another layer.
2. Model-assisted findings are probabilistic and fallible
A model may misunderstand a prompt, overlook relevant language, rely on an incomplete reference, infer an unsupported fact, assign the wrong severity, produce an inconsistent explanation, or fail to return valid structured output. The same underlying behavior may be judged differently after a model, evaluator prompt, rubric, verifier, or provider change.
Accordingly, model-assisted evaluation can produce both false positives—flagging acceptable behavior—and false negatives—missing problematic behavior. Confidence or verification metadata communicates part of the evaluation state; it is not a probability of truth or a guarantee of correctness.
High-severity verification is an escalation layer, not an independent audit or certainty mechanism. A confirmed critical finding can still be wrong or incomplete, and an overturned finding does not establish that every risk has been excluded.
3. Results are bounded by inputs, coverage, and observable behavior
Results depend materially on:
- the prompts, cases, assertions, rubrics, severity labels, and reference facts supplied;
- which knowledge sources are selected, current, complete, and internally consistent;
- the request template, response selector, credentials, timeout, and endpoint behavior;
- the number and diversity of test cases and the scenarios the customer chose to omit;
- the sampled response returned during that particular execution;
- the configured evaluator, verifier, prompt, rubric, and scoring versions; and
- external model, network, customer-agent, and provider availability and behavior.
Black-box API testing observes only the configured requests and selected responses. It does not inspect customer source code, internal model reasoning, every retrieval result, every user conversation, every tool call, or every future response unless that information is expressly included in supported test inputs.
A passing test supports only the behavior covered by that test at that time. It does not establish complete product correctness, comprehensive policy adherence, or absence of untested failure modes.
4. How to interpret scores, verdicts, and errors
Evaluation states distinguish pass, warning, failure, critical failure, evaluator error, and agent execution error. An evaluator error means the evaluation could not produce trusted final evidence. An agent execution error means the configured request did not produce a usable response. Neither is a pass.
A health score is a versioned aggregate of available evidence. Weighting and caps are designed to keep warnings, failures, critical findings, and execution/evaluator errors visible, but any aggregate necessarily simplifies the underlying cases. Review the component findings, execution evidence, deterministic outcomes, and verification state—not only the number.
A baseline comparison can identify defined changes between comparable case versions. A stable comparison means the configured comparison rules did not identify a covered regression; it does not guarantee unchanged behavior outside those rules. Changed case sets may be incomparable rather than safely inferred.
5. Agents, models, providers, and results change over time
Customer agents can vary between executions because of model sampling, prompts, retrieval, tools, user state, deployment changes, provider changes, or data updates. Evaluator and verifier models can also change. Network conditions, timeouts, and external outages may affect execution evidence independently of answer quality.
A prior pass does not predict every future response, and a prior failure does not prove every future response will fail. Re-run important cases after material changes and use appropriately reviewed baselines, test diversity, and independent monitoring.
Stored results identify relevant evaluator, prompt, rubric, scoring, and comparison versions where available. Version records improve traceability; they do not eliminate model variability or make results reproducible across a changed external provider.
6. Human review remains required
Customers should:
- review critical, surprising, ambiguous, or high-impact findings before acting;
- inspect the underlying test definition and permitted evidence, not only the headline;
- confirm important facts against authoritative and current source material;
- repeat or independently test consequential findings where appropriate;
- keep test cases, knowledge, rubrics, and baselines current and representative;
- investigate evaluator and agent-execution errors rather than treating missing evidence as success; and
- maintain release, rollback, incident, and customer-support processes outside the Service.
NoirGen Agent QA assists reviewers; it does not assume responsibility for the customer’s agent, source material, release decision, end-user communication, or response to a finding.
7. No professional advice, certification, or compliance determination
Results are not legal, regulatory, medical, financial, employment, credit, insurance, security, accessibility, fairness, or other professional advice. They are not a SOC 2 report, penetration test, HIPAA determination, safety certification, audit opinion, or government approval.
Do not use Service output as the sole basis for decisions affecting a person’s rights, safety, eligibility, access, or material opportunities. Obtain qualified domain review and use legally required validation, notices, explanations, appeal rights, and oversight.
Marketing descriptions such as “detect regressions” or “catch bad answers” describe the Service’s intended quality-assurance function. They do not promise detection of every hallucination, policy issue, unsafe response, outage, or compliance problem.
8. Questions, suspected errors, and feedback
If a result appears wrong, preserve the run and case identifiers, note the expected behavior, and review the authorized evidence available in the Service. Do not send endpoint credentials, private keys, payment-card data, or unnecessary Customer Content by email.
Product questions and suspected evaluation defects may be sent to [email protected] or started through the Support page. No response-time or resolution guarantee is stated by this pre-launch draft.