aíEmail

CALYR.AI · White Paper 01 · September 2026

From Black Box to Scientific Evidence

A framework for governable AI in engineering and medicine.

Fast models are useful only when the evidence that makes them trustworthy remains attached.

Executive summary

Enterprise AI governance has largely begun with visibility: which systems are in use, who has access, what data enters them, what they cost, and whether their use is permitted. This is necessary, but insufficient for scientific and engineering AI.

A system can be fully inventoried, access-controlled and policy-compliant and still produce a scientifically invalid result.

Scientific AI therefore needs a second governance layer: evidence governance. The central question is not only whether an AI system can be used. It is also what evidence supports a result, under which assumptions, inside which validity domain, with which uncertainty, and what should happen when those conditions fail.

CALYR.AI proposes an architecture in which models, surrogates, simulations, experiments, assumptions, uncertainty estimates, validation results and human decisions are connected as an explicit Scientific AI Evidence Graph.

Evidence · Surrogate · Oracle · Verification · Decision

01 · From Shadow AI to Shadow Reasoning

Knowing the model is not the same as knowing the evidence.

Enterprise governance has a visibility problem. AI tools can be adopted faster than organisations can discover, approve and monitor them. Scientific AI introduces another problem: an organisation may know exactly which model generated a result and still be unable to reconstruct the scientific justification behind it.

Which physical assumptions were active?

Was the query inside the model's validation domain?

What uncertainty belongs to the prediction?

Which reference simulation, experiment or measurement validates it?

Who accepted the output for the downstream decision?

We call this shadow reasoning: a result whose computational origin is visible but whose scientific justification is not.

02 · The governance gap

The governed object must be the result, not only the software.

A scientific result depends on a chain: physical or biological system, measurement or simulation data, mathematical representation, learned or reduced model, inference request, uncertainty model, validation evidence, decision rule and human owner. A failure can arise at any link.

A surrogate may be accurate globally but unreliable in a local region. A high-fidelity solver may itself use assumptions that are invalid for a new regime. A patient-specific geometry may be segmented correctly while boundary conditions remain weakly constrained.

Governance therefore has to operate at the result level.

03 · Evidence, not explanation alone

An explanation is not a validation.

A scientifically governed result should carry five evidence dimensions.

Provenance
Model version, training data or simulation campaign, preprocessing, solver versions, parameter sets and transformations.
Validity domain
The conditions under which the model has been shown to work.
Uncertainty
The uncertainty attached to the prediction, including data, approximation, measurement and model-form uncertainty where possible.
Verification and validation
Reference simulations, analytical limits, experiments, held-out cases, prospective measurements and cross-method comparisons.
Decision ownership
The person or governed process allowed to act on the result.

04 · Scientific AI Evidence Graph

Evidence should be traversable.

CALYR.AI represents the scientific chain as a graph rather than as disconnected documentation. Nodes may include physical systems, geometry, datasets, experiments, simulations, solvers, models, surrogates, assumptions, parameters, inferences, uncertainty estimates, validation results, decisions and versions.

Relations such as derived from, trained on, calibrated against, assumes, valid within, evaluated by, contradicts, supersedes, escalated to and approved by connect those objects.

What produced me? What supports me? Where am I valid? What could invalidate me? What happens next?

05 · Governed surrogates

A surrogate should behave like a scientific instrument.

A governed surrogate should carry an immutable model version, reference model or experimental source, training and validation domains, error distributions, uncertainty estimates, out-of-distribution criteria, known failure modes, retraining history and intended decision context.

The most important property is not raw speed. It is conditional trust. The surrogate should answer quickly when evidence supports the query and decline, escalate or request new evidence when it does not.

06 · Oracle

Acceptance is a separate computational task.

In the CALYR.AI architecture, the Oracle is not a second predictive model claiming to know the truth. It is the decision layer that evaluates whether the available evidence is sufficient to accept a surrogate result.

Surrogate prediction + uncertainty + validity + evidence state = acceptance decision

The Oracle can accept, accept with qualification, request verification, request new data, escalate to high-fidelity simulation, or reject a case as outside the supported domain.

07 · Active verification

Governance can choose the next experiment.

Scientific AI requires a continuous view of validation. When a model encounters a consequential region of uncertainty, the system should identify the next most valuable verification action: an additional CFD simulation, experiment, structural reconstruction, boundary-condition measurement, clinical measurement or targeted design point.

This connects governance to active learning and design of experiments. The question becomes: which new piece of evidence would reduce the uncertainty that matters for the decision?

08 · Human responsibility

Human-in-the-loop must be operational.

A reviewer should see what is being predicted, the uncertainty, whether the case lies inside the supported validity domain, which evidence supports the result, which assumptions dominate, whether contradictory evidence exists and which escalation options are available.

09 · Reference architecture

Evidence · Surrogate · Oracle · Verification · Decision

Evidence
Measurements, experiments, reference simulations, constraints and provenance.
Surrogate
Fast reduced or learned models that approximate an expensive physical or computational process.
Oracle
Acceptance logic using uncertainty, domain validity, evidence quality and decision context.
Verification
High-fidelity simulation, experiment or measurement invoked when the evidence state is insufficient.
Decision
A human or governed downstream process uses the result together with its evidence state.

The architecture is cyclic. Verification generates new evidence. New evidence can recalibrate the surrogate, update uncertainty, modify the validity domain and alter future Oracle decisions.

10 · Patient-specific aortic modelling

Fast inference without detaching from physiology.

A patient-specific workflow can reconstruct anatomy from imaging, define boundary conditions, execute reference CFD or particle-based simulations, train or adapt a fast surrogate, estimate uncertainty, detect whether the query lies inside the supported domain, compare against measurement or experimental evidence, escalate uncertain cases, and present the result with its evidence state for expert review.

The surrogate does not replace clinical judgement or high-fidelity modelling. It allocates expensive evidence where it is most informative while routine exploration becomes fast.

11 · Engineering design

High-fidelity simulation becomes selective evidence.

The same architecture applies to structural and multiphysics design. A governed surrogate can explore geometry, loading or material parameters rapidly while the Oracle tracks where approximation remains trustworthy. Selected states are verified using the reference solver, and the new evidence updates the model.

12 · Regulatory and assurance context

Evidence architecture complements formal compliance.

The EU AI Act applies progressively. Article 50 transparency obligations apply from 2 August 2026. European Commission guidance for high-risk systems emphasises risk management, data quality, documentation and traceability, human oversight, accuracy, cybersecurity and robustness, with the relevant high-risk provisions applying according to system category.

NIST's AI Risk Management Framework and AI Resource Center emphasise testing, evaluation, verification and validation. In 2026, NIST published a draft TEVV-Athlon framework for context-specific AI assessment.

For medical devices, the U.S. FDA describes AI-enabled devices through a total-product-lifecycle perspective spanning development, validation, deployment, monitoring, maintenance and modification.

13 · Minimum evidence contract

Attach evidence to every prediction.

prediction
model_version
input_signature
validity_status
uncertainty
reference_evidence
verification_status
decision_owner
timestamp

Mature systems can add assumption sets, data and solver lineage, calibration state, out-of-distribution score, counter-evidence, change history and approval records.

14 · CALYR.AI

Fast models. Evidence attached.

CALYR.AI develops scientific AI architectures for complex biophysical and engineering systems. The platform combines a Surrogate Engine for fast physics-aware prediction, Scientific AI for model-aware inference and uncertainty, an Oracle for acceptance and escalation, an Evidence Graph for provenance and validation, and interfaces for integration into research and engineering workflows.

Conclusion

Every consequential prediction should carry its evidence with it.

The next governance challenge is not only discovering which models organisations use. It is determining whether the reasoning behind a consequential computational result remains supported.

A trustworthy scientific AI system should expose where a prediction came from, where it is valid, how uncertain it is, which evidence supports it, what would trigger verification and who owns the resulting decision.

References

  1. Xensam. Out of the Shadows: Governing Enterprise AI.
  2. European Commission. AI Act.
  3. European Commission. Transparency obligations under Article 50.
  4. NIST. AI Resource Center.
  5. NIST. TEVV-Athlon Framework.
  6. U.S. FDA. Artificial Intelligence-Enabled Medical Devices.