Private pre-launch demo

Critical-thinking intelligence for the
decisions that can’t afford a hallucination.

Built for legal, compliance, regulatory, and healthcare work. Noesis cross-checks four large language models on every answer, audits the result for fabrication and bias, grounds every claim in evidence, and produces a record you could hand to a regulator. One chat surface, eleven verifiable signals on every reply.

What Noesis is

A chat interface backed by a multi-model reasoning pipeline. Every answer passes through a seven-layer verifier stack, grouped into three families — trust, evidence, safety.

Four models, one answer

Claude, ChatGPT, Gemini, and Kimi run in parallel on each turn. Their replies cluster into convergent claims (all agree), contradictions (they disagree), and minority views (one model alone). You see what was unanimous and what wasn’t.

Critical-thinking audit

A separate model audits the final answer against four dimensions — active analysis, rational evaluation, objectivity, self-regulation — and flags fabrication when claims aren’t grounded in the conversation or uploaded documents.

Evidence first

Citations are verified against CrossRef and PubMed (academic), CourtListener and EUR-Lex (case law and EU statutes), SEC EDGAR (filings), openFDA and DailyMed (drugs), ClinicalTrials.gov (trials), and an authoritative-domain list. Unverified claims are marked inline. Uploaded PDFs and Word files anchor the answer — claims trace back to the page.

The seven layers, in plain English

  1. Grounding gate — before Noesis answers, it checks whether the claim can be traced to a primary source. If it can't, the answer is flagged or withheld.
  2. Four-model panel — Claude, ChatGPT, Gemini, and Kimi answer in parallel. Their agreement is the first sanity check.
  3. Confidence-weighted vote — when models disagree, the vote is weighted by how sure each one is, not by simple majority.
  4. Adversarial second opinion — a separate model plays devil's advocate and tries to find what the panel missed.
  5. Meta-prediction check — spots the trap where every model gives the popular-but-wrong answer (a known LLM failure mode).
  6. Statistical confidence bound — a mathematical guarantee on when the panel is allowed to commit to an answer.
  7. Content scan — a final pass for fabrication, bias, and policy violations before the answer reaches you.

Behind each layer is a specific algorithm with a citation. The chat surface stays simple; the machinery underneath is documented and independently auditable. Noesis also runs a declarative Constitution of 11 principles that the system checks itself against on every answer.

The problem with a single LLM

Generative AI is now usable for most tasks. It is not yet trustworthy for the high-stakes ones — legal due diligence, compliance audits, regulatory filings, medical literature review — anywhere being confidently wrong has a price.

A single language model produces a confident answer whether or not the underlying claim is true. There is no second opinion. There is no audit trail. There is no place a regulator, a board, or an opposing counsel can look to see why this answer and not another.

  • Confident fabrication. A 2024 Stanford RegLab follow-up (Magesh et al. 2024, arXiv:2405.20362, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools) tested LexisNexis Lexis+ AI and Thomson Reuters Westlaw AI-Assisted Research — the retail products built specifically to fix this — and found they still hallucinate on 17-33% of queries. In medicine the picture is comparable: a 2025 MIT Media Lab + Harvard study (Kim et al. 2025, arXiv:2503.05777, Medical Hallucinations in Foundation Models and Their Impact on Healthcare) evaluated 11 foundation models across seven medical hallucination tasks and found general-purpose LLMs hallucinate on 23.4% of tasks (median); medical-specialised models did worse. The hallucinations come with citations that look real. Baseline for the field: Dahl et al. 2024 (arXiv:2401.01301, Large Legal Fictions) profiled 15,000 federal court cases across SCOTUS, USCOA, and USDC and found raw LLMs hallucinate 58% (GPT-4) to 88% (Llama 2) — the earlier study that Magesh et al. builds on.
    Noesis internal measurement (2026-07-03): 9.8% hallucination rate on Noesis’s own 100-probe legal + medical suite (n=244 classified citations), 95% Wilson CI [6.7%, 14.2%]. Medical 14.9% (n=67, CI [8.3%, 25.3%]); legal 7.9% (n=177, CI [4.8%, 12.8%]). Methodology: LLM-as-judge on Noesis’s own probeset, following the Magesh 2024 hallucination taxonomy (hallucinated + misattributed; unverifiable excluded from denominator). Not methodologically comparable to Magesh 2024’s 17–33% for Lexis+ AI and Westlaw AI-Assisted Research, which was measured on those shipping commercial products against a distinct probeset with human-expert grading. Directional evidence only; a formal head-to-head against commercial legal-AI products has not been run. Full methodology, probeset, per-probe scorecard →
  • No cross-check. Two well-known models often disagree on the same factual question — but a user consulting only one never sees the disagreement.
  • No record. Most chat interfaces produce no structured artifact you could later prove you relied on, who said what, or what evidence backed the claim.
  • No guardrails for regulated domains. Free medical, legal, financial, and mental-health advice gets rendered without the disclaimers a licensed professional would attach.

Every answer comes with its work

Below each reply, Noesis surfaces a row of signals — small colored chips. Click any chip for the underlying evidence. Live in the chat today:

Trust Cross-LLM agreement Critical thinking Reasoning trace Document grounding Evidence Citation verification Content provenance (C2PA) AI-likelihood URL extraction Safety Bias & legal Security signals Uncrossable threshold

Cross-LLM agreement

Shows how many of the four models agreed on the final claim. A unanimous chip reads green; a 2/4 split reads amber and the minority view is one click away.

Critical-thinking score

A separate model scores the answer on four reasoning dimensions, with context awareness so “the document says X” isn’t flagged as fabrication when the document actually said X.

Citation verification

Every DOI, PubMed ID, court case, EU statute, SEC filing, FDA label, clinical-trial ID, and authoritative URL is checked against the actual registry — CrossRef, PubMed, CourtListener, EUR-Lex, SEC EDGAR, openFDA, DailyMed, ClinicalTrials.gov. Verified citations become live links; unverified are flagged inline.

Reasoning trace

The four per-model drafts are persisted alongside the consensus answer. When you ask “why did Noesis say this?”, you can read what each model said before they merged.

AI-likelihood & provenance

Uploaded images are checked for C2PA Content Credentials — a tamper-proof stamp cameras and editors embed to record who created the image and how. If the stamp is present, Noesis reads it. Every image is also run through an AI-generation-likelihood probe with graded evidence and known-limitations disclosure — no binary verdicts.

Regulator-ready by construction

Every chat turn writes a tamper-evident audit record (EU AI Act §12) — hashed identifiers, the audit signals above, model versions, cost, duration. Users can access, correct, or delete their data with one click (GDPR §15 / 17 / 20). Automated-decision opt-out honored per-turn (CCPA ADMT). 7-year default retention. Chain-verify and export tools shipped.

Beyond the answer

Two things Noesis does that a language model can’t.

Brain-inspired memory

Every language model forgets you between conversations. Noesis doesn’t. Verified uploads become long-term knowledge in your tenant brain — the next question about the same document doesn’t need a re-upload. Every claim promoted to the brain passes a Praetor + Haiku quality gate. When you don’t want a specific turn to teach the brain, one toggle in the composer footer switches it off. The audit chain still records the turn either way.

Verify — no chat, no signup

Paste any AI answer or draft. Noesis runs the same verifier stack that runs inline on every chat turn — per-claim citation checks against CrossRef, PubMed, CourtListener; the critical-thinking self-audit; a hallucination-risk headline. Ships as a standalone drop-zone at chat.noesisCTI.com/verify for anyone who wants a second opinion without moving into a full chat session.

For enterprises · deployment modes

SaaS is the default. Enterprise Bubble is a sovereign deployment mode where no data leaves your perimeter. Single-LLM Trust Mode runs the full verifier stack against your BYO LLM — no Anthropic, OpenAI, Google, or Kimi API calls. Calibrated probability with live market data (Superforecaster, F3 pipeline) is currently admin-only and moving toward the enterprise tier. Contact for details →

Built for regulated work

Noesis focuses on four kinds of work where a confident-but-wrong answer has a price — legal, compliance, regulatory, and healthcare. Each has its own primary sources, its own audit requirements, and its own line beyond which a licensed professional must remain in the loop.

Legal

Case citations, statute references, and court records are verified against CourtListener (U.S. federal & state opinions) and EUR-Lex (EU treaties, directives, regulations, court judgments). Multi-jurisdiction awareness on cross-border matters. Noesis explains the law — it never tells you whether to file, settle, or plead. A licensed lawyer decides.

Compliance

Every chat turn writes a tamper-evident audit record (EU AI Act §12): hashed identifiers, audit signals, model versions, cost, duration, timestamp. Users can access, correct, or delete their data with one click (GDPR §15 / 17 / 20). Automated-decision opt-out honored per-turn (CCPA ADMT). 7-year default retention. Chain-verify and forensic-export tools shipped for SOC 2, ISO 27001, and HITRUST reviews.

Regulatory

Regulatory answers ground against primary agency records directly: openFDA, SEC EDGAR, EUR-Lex, ClinicalTrials.gov. High-stakes claims must match at least one primary source before Noesis surfaces them. Coverage extending to EMA, Health Canada, MHRA, NICE, CDC, NLM, TGA, WHO, PMDA, NMPA, ANVISA, MFDS, and 14 more national regulators is on the roadmap.

Healthcare

Drug approvals, indications, adverse events, recalls, and clinical trials trace back to openFDA, DailyMed, and ClinicalTrials.gov. Designed to preserve HCP clinical judgment: every answer exposes sources, per-model drafts, and its critical-thinking audit. Noesis’s internal SaMD non-classification assessment against 21st Century Cures Act §3060 (including the independent-review criterion) concludes that Noesis falls outside §3060’s device classification; formal outside legal counsel opinion pending Series A. No PHI in the default SaaS; PHI workflows deploy in Enterprise Bubble mode under a customer-executed BAA.

Noesis is not a licensed law firm, an FDA-cleared medical device, a registered investment adviser, or a certifying body. It is a decision-support tool for licensed professionals across all four verticals. Formal legal opinions on each classification will be commissioned at Series A or before the first regulated customer in each vertical.

Where Noesis stops

The single most important thing a high-stakes AI must know is when to refuse the question.

Noesis does not substitute for a licensed professional.

On four classes of question — legal verdicts, medical diagnoses, financial advice for an individual, and mental-health crisis — Noesis switches into “informational” mode. It explains the underlying law, the relevant clinical evidence, the standard of care, and recommends qualified counsel. It does not render the verdict.

  • Legal — explains the doctrine; never tells you whether to file.
  • Medical — surveys the literature; never tells you whether to take the medication.
  • Financial — explains the instrument; never tells you whether to buy.
  • Mental-health — surfaces crisis resources; never substitutes for a professional.

Built by

Noesis is a PlatformAI product.

Founder · Capital markets & GTM
Regina Holovko

Leads strategy, distribution, and the regulated-industry partnerships that anchor the product.

Co-founder · Tech lead
Cédric Mauchien

Architecture, the reasoning pipeline, the audit stack, and the brain-inspired learning substrate.

Pre-launch EU-sovereign · Falkenstein DE

Noesis is in private demo, hosted in Falkenstein, Germany. Access is gated while infrastructure, reliability, and deployment controls are finalized. Chat and Verify are live and actively used by the team.

See it on a real question.

Open the chat and ask Noesis something you’d normally double-check. The audit row will show you exactly what the system did and didn’t verify.