AI claims & evidence register
AI claims become more useful when their boundaries are visible. This register records a specific statement, the primary document behind it, what that document establishes and what would require further evidence. It is a source-reading resource, not a league table or an independent benchmark laboratory.
By Nicholas Thackray. First edition: 16 September 2026. Three entries; all sources checked on that date. Documentary review only. No model testing or independent experimental replication is claimed.
How to read an entry
A company announcement is strong evidence of what the company announced. It is a different kind of evidence from a reproduced experiment or a customer's measured result. Each entry therefore separates the source's statement from Nuvastra's interpretation. A finding marked “documented” concerns the stated documentary fact; it does not certify every product associated with the organisation.
Use the record ID when referring to an entry. The source date belongs to the original publication; the checked date records our inspection. An undated or changing source is identified as such. Historical findings remain tied to their original versions and conditions.
NVE-001: what Claude’s constitution establishes
Subject: Anthropic's account of intended Claude behaviour. Evidence type: first-party policy and training description. Source date: the reviewed page does not display a publication date. Checked: 16 September 2026. Finding: documented intentions; universal compliance is not established.
Source observation. Claude's constitution describes Anthropic's intended values and behaviour for Claude and says the document influences training. Its introduction explicitly acknowledges that actual behaviour can diverge from those intentions. The page also distinguishes its mainline, general-access models from some specialised models.
Nuvastra's interpretation. The document supports a statement about training direction. Reading it as proof that every Claude response follows every principle would go beyond the source. This is a distinction about evidence, not a finding that a particular deployment has passed or failed a safety test.
What further evidence would help? For a concrete application, record the model version, tools, instructions and allowed actions. Define observable success and failure before running the assessment. A hypothetical document assistant, for example, could be assessed on whether it identifies missing evidence and avoids inventing a source. Report both the successful and unsuccessful cases, with the test conditions. That example is a proposed method, not a result measured by Nuvastra.
Related reading: Anthropic company profile.
NVE-002: reading Mistral 7B’s historical benchmark claim
Subject: the original Mistral 7B release. Evidence type: developer-reported evaluation. Source and event date: 27 September 2023. Checked: 16 September 2026. Finding: a historical vendor comparison, not a current universal ranking.
Source observation. Mistral's launch announcement compares Mistral 7B with Llama models using evaluations rerun by Mistral. Its summary uses broad language about outperforming Llama 2 13B; a later figure caption qualifies the knowledge benchmarks as being on a par. The page specifies different shot counts and notes differences from the Llama 2 paper's evaluation, including its handling of MBPP and TriviaQA.
Nuvastra's interpretation. The detailed conditions and qualification belong alongside the headline comparison. The source does not establish that Mistral wins on every possible task, every later model version or a reader's production workload. Nuvastra has not rerun these experiments.
What further evidence would help? Preserve the exact checkpoints, datasets, prompts, scoring rules and runtime configuration before attempting replication. A business comparison would additionally need representative tasks and an agreed quality threshold. If one candidate produces an acceptable answer quickly but another needs several retries, a useful comparison records the complete cost and elapsed time per accepted outcome. This is a suggested measurement design, not a claim about these two models' actual operating costs.
Related reading: Mistral AI company profile.
NVE-003: what Google DeepMind’s 2023 date means
Subject: formation of the combined Google DeepMind unit. Evidence type: organisational announcement. Source and event date: 20 April 2023. Checked: 16 September 2026. Finding: documented formation announcement; distinct from the earlier history of the constituent teams.
Source observation. Google's 20 April 2023 announcement describes a new unit bringing together DeepMind and the Brain team from Google Research. It names that combined group Google DeepMind. The announcement also refers to the teams' earlier work, making clear that their research histories predate the combination.
Nuvastra's interpretation. A database field labelled “founded” is ambiguous unless it names the entity and event. The date of this combination should not silently become the founding date of every predecessor organisation, the release date of Gemini or the date of an AlphaFold result.
How to record it. Keep separate fields for entity name, event type, event date, source and checked date. In this record, the event type is the announced formation of the combined unit. If a future entry covers an acquisition or product launch, retain its own evidence rather than reusing this date. That simple distinction makes timelines easier to audit and reduces errors when readers combine information from several sources.
Related reading: Google DeepMind company profile.
Scope, corrections and revision history
This first edition covers three deliberately different kinds of evidence: intended behaviour, a reported experiment and an organisational event. It is not a comprehensive audit of these companies. Selection does not imply endorsement, wrongdoing or a comparative safety rating.
To suggest a correction, use Nuvastra's contact page and include the record ID, the statement concerned and supporting evidence. Changes should identify what was amended and why; a checked date should change only after another source review.
Revision 1 — 16 September 2026: register launched with NVE-001, NVE-002 and NVE-003. No independent tests performed.
Continue with the AI company directory, AI Fundamentals or the AI glossary.
