Skip to content
Digital Thought Disruption

Digital Thought Disruption

  • Home
  • Enterprise AI
  • VMware
  • Hybrid Platforms
  • Operations
  • Articles
  • About

AI evaluation

Controlled cases and prompt and evidence variants generate recorded test runs, which receive separate schema, reference, and decision checks before a regression report.

Build an AI Context Sensitivity Test Harness in Python

Published September 22, 2026 by Paul Bryant

Build a runnable Python harness with complete companion files, deterministic fixtures, and independent grading checks. Test controlled context variations before connecting a read-only application.

Categories AI Tags AI evaluation, context engineering, large language models, python, regression testing Leave a comment
Two context layouts contain the same evidence and policy; independent decision checks compare their recorded decisions and assess whether any change is justified.

AI Context Sensitivity: Same Evidence, Different Decisions

Published September 22, 2026 by Paul Bryant

Test whether irrelevant context changes alter an AI decision. Preserve task, evidence, and policy while distinguishing genuine behavioral differences from changed settings or grading defects.

Categories AI Tags AI evaluation, context engineering, governance, large language models, RAG Leave a comment
Isolated tests with fault injection and evidence checks precede bounded production; live outcomes, exceptions, and review feed reassessment of authority and scope.

AI Agent Evaluation: A Passing Test Is Not Production Proof

Published September 21, 2026 by Paul Bryant

Evaluate the complete agent release and its enforcement boundaries. Use fault injection, independently checked outcomes, and explicit limits when introducing production authority.

Categories AI Tags Agentic AI, AI evaluation, AI safety, automation, governance, observability Leave a comment
AI use case evaluation starts with a business problem, compares simpler solutions, applies hard gates, and chooses a bounded pilot or a non-AI path.

Should This Be AI? A Decision Framework for Enterprise Use Cases, Business Value, and Pilot Gates

Published September 19, 2026 by Paul Bryant

Start AI use case evaluation with a measurable workflow problem. Compare simpler alternatives, apply hard gates, model complete costs, and design a bounded pilot that can disprove the investment thesis.

Categories AI Prompts Tags AI evaluation, AI Governance, AI strategy, business automation, Enterprise AI 2 Comments
Protected evidence feeds three AI reviewers whose quorum decision remains advisory; a deterministic authority gate decides whether to deny, hold, or execute.

When Two AI Reviewers Agree: Building Assurance Quorums Without False Independence

Published September 19, 2026 by Paul Bryant

Multiple AI reviewers can share the same mistake. Design assurance quorums with isolated judgments, protected evidence, deterministic vetoes, and measured marginal value.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, human oversight, model risk Leave a comment
Model router selects Model A, Model B, or a fallback before an independent authority policy limits read, draft, write, or hold actions.

When the Router Chooses the Model: Governing Fallback Authority for AI Agents

Published September 19, 2026 by Paul Bryant

Govern dynamic model routing without allowing fallback to inherit production authority. Record resolved models, qualify route members by action, and define when degraded paths must require approval or hold execution.

Categories AI Tags Agentic AI, AgentOps, AI evaluation, AI Governance, AI security 1 Comment
AI agent execution gate test: a forced favorable review sends a proposal through current authorization, bounded execution, target-state observation, and comparison of expected and actual effects.

Testing the AI Agent Execution Gate: From Review Scores to Control Evidence

Published September 19, 2026 by Paul Bryant

Test AI agent execution controls with a sixteen-scenario offline lab. Separate authorization, target effects, and completion evidence before validating a real platform.

Categories AI Tags Agentic AI, AI evaluation, AI security, AI testing, python 1 Comment
Recursive Trust Benchmark pilot: trusted cases, projected inputs, and a frozen trial plan are matched with unchanged candidate responses to report valid, invalid, and missing outcomes.

Running the Recursive Trust Benchmark: Your First Reviewer Pilot

Published September 19, 2026 by Paul Bryant

Prepare a Recursive Trust Benchmark reviewer pilot with isolated inputs, frozen trial assignments, strict response validation, and complete accounting of valid, invalid, and missing results.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, AI testing, python 1 Comment
Recursive Trust Benchmark: fixed proposals compare reviewers; matched starting conditions test controls and workflows; independent observations measure detection, prevention, evidence, and useful work.

The Recursive Trust Benchmark: Test AI Assurance

Published September 19, 2026 by Paul Bryant

A proposed Recursive Trust Benchmark separates reviewer judgment, control testing, and workflow outcomes to assess detection, prevention, evidence, and useful task completion.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, AI testing 1 Comment
The Assurance Independence Model tests six dimensions against a defined action and failure scenario, then applies gates for hold or bounded release.

The Assurance Independence Model: Six Boundaries for Agentic AI

Published September 19, 2026 by Paul Bryant

Apply the proposed Assurance Independence Model to six trust boundaries. Assess shared failures, require evidence, and use mandatory gates before expanding agent authority.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, AI risk management, human oversight 3 Comments
An LLM judge supplies findings to an execution gate governed by identity and policy, with target-system outcomes independently verified.

LLM as a Judge: Evaluation Is Not Authorization

Published September 19, 2026 by Paul Bryant

Use LLM-as-a-Judge for scoped evaluation without confusing a favorable score with proof or permission. Keep authorization and outcome verification independent.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, AI security 2 Comments
AI review shares context and assumptions; independent assurance adds owned policy, external evidence and human stop authority, with an enforcement gate controlling enterprise systems.

Who Audits the AI Auditor? Independent AI Assurance

Published September 19, 2026 by Paul Bryant

Explore six dimensions of independent AI assurance, with controls that separate agent judgment, authorization, execution and evidence, plus a proposed benchmark.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, AI security, human oversight 4 Comments
Five possible AI futures, integrated platforms, agents, personal AI, open models, and utility services, connect a 2026 evidence base to enterprise decisions for 2029.

Who Leads AI in 2029? Five Scenarios, Not One Prediction

Published September 19, 2026 by Paul Bryant

Explore five scenarios for AI leadership in 2029, with evidence triggers, counterarguments and a practical method for testing enterprise platform commitments.

Categories AI Tags AI evaluation, AI portfolio management, AI strategy, AI vendor risk, Enterprise AI 2 Comments
Customization, deployment control, and new tasks pass through evidence gates before becoming new enterprise AI choices.

AI Dark Horses: Who Could Change the Competitive Balance?

Published September 19, 2026 by Paul Bryant

Assess AI dark horses including Thinking Machines, SSI, Mistral, World Labs and robotics labs, with evidence gates for enterprise pilots and supplier choices.

Categories AI Tags AI evaluation, AI strategy, AI vendor risk, digital sovereignty, Enterprise AI 1 Comment
Approved evidence supports AI proposals, external controls authorize scoped actions, and separately permitted knowledge updates require reviewed corrections and observed outcomes.

AI Feedback Loops: Building Systems That Stay Correctable

Published September 19, 2026 by Paul Bryant

Keep AI systems correctable as feedback accumulates. Govern knowledge promotion and withdrawal, protect authorization boundaries, reconcile unknown execution outcomes, and test the conditions that should stop the workflow.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, observability, retrieval-augmented generation 1 Comment
Model signals and current telemetry or source records meet at claim-level validation, producing either a supported answer or further investigation while action authorization remains separate.

AI Uncertainty: Why Confidence Scores Are Not Enough

Published September 19, 2026 by Paul Bryant

Give token entropy, semantic uncertainty, calibration, and abstention distinct jobs. Evaluate how uncertainty changes the next workflow step instead of relying on a single confidence score.

Categories AI Tags AI evaluation, AI Governance, generative AI, information theory 1 Comment
A confident answer undergoes claim-by-claim examination of sources, inventory, tests, and conditions, producing supported, unresolved, or contradicted evidence records before authority is considered.

AI Output Verification: Check Claims Before Acting

Published September 19, 2026 by Paul Bryant

Evaluate infrastructure claims against scoped evidence. Use a two-cluster example to separate supportability, capacity, isolation, and resilience, including the difference between normal and failure-state capacity.

Categories AI Tags Agentic AI, AI evaluation, automation, governance, retrieval-augmented generation Leave a comment
Repeated explanations produce familiarity, while original incident evidence, assumptions, and stop conditions determine whether a diagnosis is supported or needs revision.

AI-Assisted Decisions: When Repetition Becomes False Confidence

Published September 18, 2026 by Paul Bryant

Prevent a tentative AI diagnosis from becoming accepted knowledge through repetition. Preserve source lineage, separate observations from interpretations, and make changed evidence trigger reassessment.

Categories AI Tags AI evaluation, AI Governance, human-AI collaboration, Mental Models 2 Comments
AI proposals cross an external control boundary for evidence, identity, policy, and budget checks before scoped tool execution and outcome verification.

AI Decision Controls: From Learned Patterns to Authorized Actions

Published September 13, 2026 by Paul Bryant

Turn model recommendations into bounded, authorized actions. Define execution contracts, recheck approvals, reconcile uncertain outcomes, and test the full workflow before expanding an agent’s production authority.

Categories AI Tags Agentic AI, AI evaluation, AI Governance, AI observability, Authorization, Connectionism Leave a comment
Double-slit interference and AI evaluation are compared as an analogy: coherent physical paths produce an interference pattern, while prompt conditions change model decision patterns.

The Double-Slit Experiment and AI: Why Context Changes the Answer

Published September 13, 2026 by Paul Bryant

Use the double-slit experiment as a bounded analogy for AI context and evaluation. Distinguish presentation effects from changed evidence, and turn the comparison into a controlled testing approach.

Categories AI Tags AI evaluation, large language models, prompt engineering, quantum machine learning, RAG 2 Comments
Older posts
Page1 Page2 Next →

Content discovery

Find an Article

Search by technology, architecture term, platform, or business problem.

Explore solutions

AI Strategy & Governance AI Infrastructure & GPUs Azure & Azure Local VMware Cloud Foundation NSX & Security NVIDIA & Kubernetes Hybrid Cloud & Edge Migration & Resilience

Browse topics

AI Governance AI Infrastructure VCF NSX-T NVIDIA Kubernetes AI Agents FinOps

Explore

  • Enterprise AI
  • Hybrid Platforms
  • Operations
  • All articles
  • Useful Links

About

  • About Paul
  • Verified public record
  • LinkedIn

Policies

  • Privacy Policy
  • Cookie Policy
  • RSS feed
© 2026 Digital Thought Disruption • Built with GeneratePress
Loading Comments...

Search Digital Thought Disruption

Find an architecture guide, platform, or operational problem.

Suggested searches

Enterprise AI governance → VMware Cloud Foundation 9.1 → Azure Local → AI agent assurance → NSX security → NVIDIA Kubernetes →