Epistemic Stress Test — Muse Spark 1.1 validated by MarCognity-AI

I ran responses generated by the model and passed them through
MarCognity-AI’s Skeptical Agent.

Not to judge the models.
Not to compare outputs.

But to make visible what usually stays hidden: the gap between how confident a text sounds and how verifiable it actually is.

Even technically correct answers contain claims that cannot be traced back to any source.

That gap has a name: Epistemic Fracture.

Github: https://github.com/elly99-AI/MarCognity-AI.git
Zenodo: https://doi.org/10.5281/zenodo.20509721

Thanks for sharing this experiment. I re-analysed the published material at claim level using DESi and then subjected the DESi analysis itself to a rule-guided Doktores self-audit.

The purpose was not to judge the legal correctness of the Muse Spark response, but to examine the epistemic structure and traceability of the claimed validation.

The main findings were:

  • 23 curated claims were extracted, typed and anchored.
  • Only 4 of 23 had domain-admissible evidence in the supplied material.
  • 18 of 23 had no concrete evidence passage attached.
  • One genuine contradiction was found between the stated method and the printed prompt: the method says that no verification, source or stage instructions were given, while the prompt explicitly requests references, direct quotations, citation checking and six named phases.
  • Two additional issues were classified more cautiously: one as a pipeline inconsistency — VERIFIED without citable evidence — and one as an unsubstantiated claim of validator independence.
  • The final “Epistemic Boundary” conclusion appears self-sealing as presented: both validator success and validator failure are interpreted as confirmation, while no explicit falsification condition is specified.

The analysis also recognises genuine strengths in MarCognity-AI, including claim decomposition, multi-source retrieval and an explicit skeptical pass. The central problem is not the idea, but source-domain gating, provenance and the absence of concrete claim-to-evidence bindings.

The full report, including all 23 claims, the evidence classifications, the three structural conflicts, limitations and the Doktores self-audit, is available here:

I would be particularly interested in your response to the three conflicts discussed on page 2. It would also help reproducibility if the exact PubMed document — or the complete source bundle supplied to the validator — were made available.