Sub-Track 02: Safety & Hallucination Mitigation

Last Audited: 2026-08-21
NUP AI-Native Verified
ISO 42001 Cl. 8.4NIST AI RMF MEASURE-2.6
In Plain Language

The term "hallucination" is frequently used as a lazy umbrella for any AI behavior a user didn't expect. In reality, AI trust failures divide into three fundamentally different problems: the system promising more than it can deliver (an expectations problem), the interface failing to show evidence and uncertainty clearly (a UX problem), and the model stating falsehoods with unwarranted confidence (a technical grounding problem). This sub-track provides the frameworks, guardrails, and detection pipelines to master all three.

CATEGORY 04 · SUB-TRACK 02

Safety & Hallucination Mitigation: The Trust Architecture

Deconstructing the "hallucination" catch-all into actionable expectation setting, trustworthy interaction design, and technical runtime verification.

In Plain Language: Why "Hallucination" Conflates Three Different Problems

The term "hallucination" is frequently used as a lazy umbrella for any AI behavior a user didn't expect. In reality, AI trust failures divide into three fundamentally different problems: the system promising more than it can deliver (an expectations problem), the interface failing to show evidence and uncertainty clearly (a UX problem), and the model stating falsehoods with unwarranted confidence (a technical grounding problem). This sub-track provides the frameworks, guardrails, and detection pipelines to master all three.

Moving Beyond the "Hallucination" Umbrella

When an AI feature returns an unexpected result, engineering teams reflexively label it a "hallucination" and scramble to edit prompt instructions. But treating all unexpected behavior as a single flaw obscures the root causes. A user complaining that an assistant gave bad tax advice might be experiencing an expectations problem (the system was asked to do something beyond its authorized capability). A user who cannot tell which document supplied a medical finding is facing a UX legibility problem (citations are invisible). And a model that invents a fictitious regulatory statute is suffering a technical grounding failure.

Each failure class demands a distinct architectural intervention. Attempting to solve UX omissions or scope overpromises with stochastic prompt engineering creates brittle, unmaintainable systems. This sub-track organizes trust and safety across these three essential domains.

Recommended Learning Path

Start Sub-Track 02: Topic 01 — Real-Time Grounding Verification & Citation Checking

Master sentence-level claim extraction, NLI entailment scoring, and bidirectional citation chips.

Start Topic 01

The Three Pillars of AI Trust & Safety

Visualizing the structural breakdown across Expectation Framing, UX Interaction Legibility, and Technical Grounding.

Three-Domain AI Trust & Safety Failure TaxonomyArchitectural marketecture diagram decomposing the generic hallucination term into Expectations & Scope Framing, UX & Interaction Trust, and Technical Grounding failure domains.THE AI TRUST TAXONOMY: DECONSTRUCTING THE "HALLUCINATION" UMBRELLADOMAIN 1: EXPECTATIONS & SCOPESystem Promising Too MuchOverpromising reasoning capability or role scopeCOMMON SYMPTOMS:• Out-of-domain compliance reliance• Arithmetic & logic guessing in LLM• Sycophantic validation of false premisesPRIMARY REMEDIATIONS:✓ Domain boundary prompt framing✓ Deterministic calculation routing✓ Assistant peer persona calibrationDOMAIN 2: UX & LEGIBILITYOpaque Interface BoundariesFailing to communicate sourcing & uncertaintyCOMMON SYMPTOMS:• Omitted or unlinked bottom citations• Silent failure / generic error strings• No human recourse or appeal queuesPRIMARY REMEDIATIONS:✓ Interactive sentence citation chips✓ Epistemic confidence indicators✓ Supervisor override & appeal circuitDOMAIN 3: TECHNICAL GROUNDINGFabrication & Adversarial FlawsUngrounded claims, confabulation & injectionsCOMMON SYMPTOMS:• Fabrication despite clean RAG context• Invented standards / false citations• Indirect prompt injection in PDF dataPRIMARY REMEDIATIONS:✓ NLI claim entailment verification✓ Dual-LLM data isolation pattern✓ Synthetic red-teaming & Wilson stats
DOMAIN 1

Expectations & Scope Framing

Occurs when product marketing, interface framing, or unconstrained system prompts lead users to treat probabilistic language models as infallible deterministic oracles across complex math, medical, or legal decisions.

KEY SYMPTOMS:
  • Users relying on conversational LLMs for out-of-domain compliance or regulatory determinations without audit trails.
  • User frustration when the model guesses arithmetic or logic proofs rather than executing dedicated calculator tools.
  • Sycophantic agreement: the model validates false user assumptions to sound helpful rather than stating boundaries.
Topics: 01, 02Explore ➔
DOMAIN 2

UX & Interaction Trust Design

Occurs when the front-end user experience presents AI output as flat, ungrounded prose without inline citation chips, sentence-level confidence indicators, graceful refusal states, or supervisor escalation paths.

KEY SYMPTOMS:
  • Citations are omitted entirely or dumped at the bottom as dead URLs with no sentence-level correspondence.
  • Silent failure or generic "Something went wrong" errors instead of informative, calibrated self-refusals.
  • Zero user recourse: operators cannot inspect retrieved source passages or trigger human supervisor appeals.
Topics: 03, 08Explore ➔
DOMAIN 3

Technical Hallucination & Grounding

Occurs when stochastic generation produces assertions that directly contradict retrieved context, invents non-existent citations, or falls prey to adversarial prompt injection payloads embedded in external data.

KEY SYMPTOMS:
  • Fabrication despite clean context: the LLM invents plausible-sounding regulations or functions not in source text.
  • Ungrounded claims interspersed seamlessly inside an otherwise factually accurate response.
  • Indirect prompt injection: untrusted external PDF text hijacks model instructions to exfiltrate private session data.
Topics: 04, 05, 06, 07Explore ➔

Defense-in-Depth: Relationship with Sub-Track 01 (RAG Systems)

A common architectural misconception is assuming that once an engineering team optimizes RAG retrieval, hallucination is solved. In reality, Sub-Track 01 (RAG Systems) and Sub-Track 02 (Safety & Hallucination Mitigation) are two complementary halves of a unified defense-in-depth pipeline.

High-precision retrieval (Sub-Track A) ensures that the language model is supplied with clean, structure-preserving source context—minimizing the frequency with which the model is starved of facts. Sub-Track B is the active runtime defense layer that catches, bounds, measures, and appeals the failures that inevitably get through.

Defense-in-Depth AI Safety & Trust PipelineArchitectural diagram illustrating how Upstream RAG Context combines with Input Guardrails, Core Inference, Output NLI Entailment, Calibrated Refusal, Human Circuits, and Regulatory Audit Logging.DEFENSE-IN-DEPTH ARCHITECTURE: SUB-TRACK 01 (RAG) ➔ SUB-TRACK 02 (SAFETY)SUB-TRACK 01: UPSTREAM RAGNoise & Starvation ReducerMaximizes factual context densityAST Layout Parsing (A.4-A.6)Preserves tables & semantic headersHybrid BM25 + Vector (A.3)Eliminates semantic search blindspotsCross-Encoder Rerank (A.7)Tightens context token efficiencyUPSTREAM OUTPUT:High-Fidelity Context Pack(Minimizes hallucination opportunity)SUB-TRACK 02: ACTIVE INTERCEPTIONRuntime Defense HarnessCatches errors that bypass retrievalInput Guardrails & Injection (B.3, B.5)Dual-LLM untrusted payload quarantineNLI Claim Verification (B.1)Sentence-level context entailment scoringCalibrated Self-Refusal (B.2)Graceful "I don't know" boundary rulesMIDSTREAM GATE:Entailment & Schema Pass(Quarantines speculative hallucination)GOVERNANCE & RECOURSEContinuous VerificationTelemetry, appeals & audit evidenceHuman-in-the-Loop (B.8)Supervisor override queues & appealsProduction Telemetry (B.6)Wilson score drift & triage flywheelsISO 42001 / NIST AI RMF (B.7)Verifiable compliance audit dossiersFINAL DELIVERABLE:Regulated Production Safety(Auditable, calibrated, trustworthy)

Upstream Noise Reduction vs. Downstream Active Interception

Sub-Track A (RAG):
Sub-Track 01 (RAG Systems) maximizes retrieval precision, structural fidelity, and context density—minimizing how often the model is starved of facts.
Sub-Track B (Safety):
Sub-Track 02 (Safety & Mitigation) deploys runtime NLI entailment, NeMo guardrails, and refusal circuits to catch the errors that get through regardless.
💡 Why Both Are Mandatory: High-precision retrieval alone cannot prevent a model from misinterpreting facts or falling for prompt injections; active safety guardrails alone cannot compensate for missing context.

Retrieval Recall vs. Epistemic Uncertainty Calibration

Sub-Track A (RAG):
Evaluates whether the vector pipeline fetched the right chunks (Context Recall, Context Precision, MRR).
Sub-Track B (Safety):
Evaluates whether the model knows when evidence is missing and refuses gracefully rather than confabulating (Calibrated Uncertainty, Faithfulness).
💡 Why Both Are Mandatory: A system with 98% retrieval recall will still encounter 2% unanswerable queries where graceful refusal is the only compliant behavior.

Engineering Mechanics vs. Regulatory Audit Evidence

Sub-Track A (RAG):
Builds the indexing, chunking, reranking, and search infrastructure.
Sub-Track B (Safety):
Produces verifiable compliance dossiers, Wilson score telemetry, and human appeal workflows mandated by ISO 42001 and NIST AI RMF.
💡 Why Both Are Mandatory: Engineering sophistication without verifiable safety telemetry fails enterprise regulatory approval.

Interactive AI Trust & Safety Incident Triage

Select a common production failure symptom to inspect its root-cause mechanism, failure domain classification, and recommended engineering interventions.

Incident Classification

Fabrication of Regulatory & Medical Facts

Technical Hallucination & Grounding

Observed Symptom: The model gives a detailed, confident response citing non-existent ISO clauses or medical studies when the retrieved documents do not contain the answer.

ROOT CAUSE ANALYSIS

The model experienced epistemic uncertainty but defaulted to autoregressive text completion rather than calibrated self-refusal, hallucinating plausible-sounding citations.

RECOMMENDED INTERVENTION

Deploy real-time NLI claim verification (Topic 01) and calibrate uncertainty thresholds with graceful self-refusal directives (Topic 02).

The 8-Topic Staged Curriculum Syllabus

The topics in Sub-Track 02 are organized into 3 logical clusters, guiding engineering teams from foundational NLI verification through UX legibility to continuous adversarial red-teaming and regulatory audit dossiers.

CLUSTER 1

Foundational Trust Framing

Core runtime verification, entailment checking, and epistemic uncertainty calibration.

9 min readSafety & Retrieval Engineers

Extracting claim spans and verifying bidirectional entailment against source documents using natural language inference (NLI).

• Sentence-level claim extraction decomposes generated text into verifiable assertions.• NLI entailment scoring flags unsupported claims prior to client streaming.• Bidirectional citation mapping ties every assertion to specific retrieved chunk offsets.
Read Topic
8 min readAI Alignment & Prompt Architects

Training models and harness circuits to detect missing evidence, measure epistemic uncertainty, and refuse politely rather than hallucinate.

• Epistemic uncertainty metrics measure when retrieved context lacks sufficient evidentiary support.• Principled self-refusal heuristics replace speculative completions with polite boundary statements.• Refusal temperature calibration prevents sycophantic guessing on out-of-domain queries.
Read Topic
CLUSTER 2

UX & Interaction Design for Trust

Enforcing deterministic output boundaries, programmable guardrails, and human escalation workflows.

9 min readFull-Stack AI Engineers

Enforcing strict structural, topical, and safety constraints using programmable guardrail frameworks (NeMo, Llama Guard, Zod envelopes).

• Programmable Colang policies intercept unsafe dialogue flows before model execution.• Zod and JSON Schema envelopes guarantee type-safe deterministic downstream consumption.• Topical and moderation rails block off-brand or dangerous hallucinations at the boundary.
Read Topic
8 min readProduct Managers & Systems Architects

Designing graceful human escalation paths, supervisor override queues, and audit-logged feedback loops when AI uncertainty is high.

• Uncertainty threshold triggers route borderline inferences to supervisor review queues.• Audit-logged human corrections feed directly into offline golden test regression sets.• End-user appeal workflows maintain operator trust and regulatory compliance transparency.
Read Topic
CLUSTER 3

Technical Detection, Defense & Governance

Adversarial red-teaming, prompt injection defenses, production telemetry, and regulatory audit dossiers.

9 min readRed Team & Security Specialists

Continuous synthetic red-teaming pipelines testing system resilience against adversarial evasion, role-play exploits, and jailbreak vectors.

• Synthetic adversary LLMs generate combinatorial prompt injection and jailbreak permutations.• Automated safety regression suites run against CI/CD pull requests before model deployment.• Evaluation scoring measures safety boundary permeability across harmful intent categories.
Read Topic
9 min readSecurity Architects & MLOps

Defending against untrusted external documents embedding hidden instructions (invisible markdown, CSS tricks) to hijack model control.

• Dual-LLM architecture separates untrusted data reading from privileged action execution.• Payload sanitization strips active instruction tokens from raw document chunks.• Least-privilege tool execution prevents prompt-injected models from unauthorized side effects.
Read Topic
8 min readQA Leads & Telemetry Engineers

Tracking Wilson score confidence intervals for factual fidelity, automated production telemetry, and weekly failure triage workflows.

• Wilson score confidence intervals quantify statistical hallucination rates in live production.• Asynchronous LLM-as-a-judge telemetry monitors faithfullness drift without adding latency.• Weekly failure triage flywheels turn flagged production hallucinations into verified test fixtures.
Read Topic
10 min readAI Governance & Compliance Officers

Mapping probabilistic safety telemetry to ISO/IEC 42001, NIST AI RMF 1.0 (MAP, MEASURE, MANAGE), and FDA Good Machine Learning Practice.

• Traceable audit dossiers link safety telemetry directly to ISO 42001 Cl. 8.4 verification requirements.• NIST AI RMF GOVERN and MEASURE evidence catalogs document risk management controls.• Automated evidence export formats simplify third-party compliance and QMS certification audits.
Read Topic
Try This with AI: Triage Production AI Trust & Safety Incidents

Use this structured prompt with your AI coding assistant to classify production failure logs into the 3 safety domains and generate immediate engineering remediation steps.

Act as a Principal AI Safety & Trust Architect. Analyze our production incident log below and perform a structured safety triage: 1. FAILURE CLASSIFICATION: Categorize the incident into one of the three core domains: - Domain 1: Expectations & Scope Framing (Overpromising system capability or role ambiguity) - Domain 2: UX & Interaction Trust Design (Opaque citations, missing confidence cues, unhandled failure states) - Domain 3: Technical Hallucination & Grounding (Ungrounded assertions, NLI entailment failure, prompt injection) 2. ROOT CAUSE MECHANISM: Explain precisely why this failure occurred at the architectural, interface, or prompt level. 3. REMEDIATION ACTION PLAN: - Immediate mitigation (e.g. prompt constraint, NeMo guardrail, refusal directive, citation UI update) - Long-term architectural prevention (e.g. dual-LLM isolation, automated red-teaming, Wilson score telemetry) - Regulatory impact mapping to ISO/IEC 42001 Cl. 8.4 and NIST AI RMF MEASURE controls. [PASTE YOUR INCIDENT LOG OR UNEXPECTED MODEL OUTPUT HERE]
Previous Section
Deterministic Unified Process
Next Track
The Four Layers of LLM Engineering

Community Discussion & Feedback

Attributed peer feedback and official Netspective architecture notes.

Was this documentation helpful?(100% found this helpful • 0 ratings)

Leave Feedback or Question

○ Loading user info...
0/2000 chars

Discussion (0)

Loading discussion thread...