This source backbone collects the primary reading paths behind VS007 — Calibration Drift. It is designed as a public source map: a way to see the documentation, research, reporting, security references, product materials, and DFEI interpretation layers that informed the issue.
The backbone is not a claim ledger and not a complete bibliography. It is a curated route through the sources most relevant to the issue’s central question:
What evidence should an AI system produce before it is allowed to continue?
01 — Deployment-Like Evaluation and Runtime Evidence
OpenAI — Deployment Simulation
Link: https://openai.com/index/deployment-simulation/
Source type: Primary institutional documentation.
Why it matters: OpenAI describes evaluation work that moves closer to realistic deployment conditions before release. VS007 uses this as part of the broader shift from static benchmark thinking toward deployment-adjacent testing.
Issue connection: Deployment-like evaluation; benchmark limits; pre-release behavioral simulation.
OpenAI Alignment — Public chat data and real-world misalignment evaluation
Link: https://alignment.openai.com/validating-public-evals/
Source type: Primary institutional research note.
Why it matters: The piece discusses limitations of public evaluations and the value of testing against more realistic conversational evidence.
Issue connection: Public-eval limits; deployment evidence; why benchmark performance is not the same as runtime safety.
OpenAI API Docs — Evaluate agent workflows
Link: https://developers.openai.com/api/docs/guides/agent-evals
Source type: Product documentation.
Why it matters: The documentation shows how agent workflow evaluation can include traces, graders, datasets, and evaluation runs.
Issue connection: Agent evaluation infrastructure; trace-based workflow inspection.
02 — Agent Control and Defense-in-Depth
Google DeepMind — Securing the future of AI agents
Link: https://deepmind.google/blog/securing-the-future-of-ai-agents/
Source type: Primary institutional research roadmap.
Why it matters: DeepMind frames AI control as a defense-in-depth problem for agentic systems, including permissions, monitoring, and behavior verification.
Issue connection: Agent control architecture; safe operation; control before continuation.
Axios — Google DeepMind prepares for rogue AI agents
Link: https://www.axios.com/2026/06/18/google-deepmind-prepares-for-rogue-ai-agents
Source type: News coverage.
Why it matters: Axios provides a secondary reporting lens on DeepMind’s agent-control roadmap and its public-policy significance.
Issue connection: Public terrain; agent-control news signal.
03 — Observability, Tracing, and Telemetry
Microsoft Foundry — Observability in Generative AI
Link: https://learn.microsoft.com/en-us/azure/foundry/concepts/observability
Source type: Product documentation.
Why it matters: Microsoft describes observability for generative AI systems, including traces, tool calls, decisions, and service dependencies.
Issue connection: Runtime telemetry; production traces; visibility into system behavior.
Microsoft Foundry — Agent Service overview
Link: https://learn.microsoft.com/en-us/azure/foundry/agents/overview
Source type: Product documentation.
Why it matters: The Agent Service overview establishes the deployment context for managed AI agents.
Issue connection: Agent workflow deployment; production context.
Microsoft Foundry — Agent tracing concept
Link: https://learn.microsoft.com/en-us/azure/foundry/observability/concepts/trace-agent-concept
Source type: Product documentation.
Why it matters: This documentation describes agent-run tracing across inputs, outputs, tool use, retries, latency, and cost.
Issue connection: Trace evidence; runtime observability; continuation evidence.
Microsoft Foundry — Agent development lifecycle
Link: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/development-lifecycle
Source type: Product documentation.
Why it matters: The lifecycle material connects development, evaluation, tracing, review, and monitoring.
Issue connection: Lifecycle framing for agent evaluation and operational review.
OpenTelemetry — GenAI semantic conventions
Link: https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/
Source type: Technical specification.
Why it matters: OpenTelemetry provides vocabulary for GenAI spans, events, attributes, messages, and instrumentation.
Issue connection: Standardized telemetry; instrumentation language.
OpenTelemetry — Generative AI metrics
Link: https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-metrics/
Source type: Technical specification.
Why it matters: These conventions describe metrics for generative AI systems.
Issue connection: Measurement layer; observability infrastructure.
OpenTelemetry — Semantic conventions for agentic systems discussion
Link: https://github.com/open-telemetry/semantic-conventions-genai/issues/35
Source type: Technical discussion.
Why it matters: The discussion reflects emerging interest in agentic observability conventions.
Issue connection: Agent telemetry watchlist; emerging instrumentation language.
AgentSight — System-Level Observability for AI Agents Using eBPF
Link: https://arxiv.org/abs/2508.02736
Source type: Research paper.
Why it matters: AgentSight explores system-level observability for AI agents, including correlations between semantic actions and system behavior.
Issue connection: Observability gap; agent-system telemetry research.
04 — Security, Red Teaming, and GenAI Risk
OWASP GenAI Security Project
Link: https://genai.owasp.org/
Source type: Security reference.
Why it matters: OWASP’s GenAI Security Project collects security categories and resources for generative AI and agentic applications.
Issue connection: Security reference layer; GenAI application risk.
OWASP Top 10 for LLM Applications
Link: https://owasp.org/www-project-top-10-for-large-language-model-applications/
Source type: Security reference.
Why it matters: The OWASP Top 10 organizes common LLM application risk categories, including risks relevant to excessive agency, prompt injection, and data exposure.
Issue connection: Application-risk framing; agentic workflow risk.
PyRIT paper
Link: https://arxiv.org/abs/2410.02828
Source type: Research paper.
Why it matters: PyRIT presents an automated red-teaming framework for identifying risks in generative AI systems.
Issue connection: Red-team tooling; adversarial testing.
garak paper
Link: https://arxiv.org/abs/2406.11036
Source type: Research paper.
Why it matters: garak provides a structured approach to probing LLM vulnerabilities.
Issue connection: Scanner-based testing; adversarial evaluation.
05 — AI Security Policy and Governance Terrain
White House — Executive Order 14409
Source type: Government policy document.
Why it matters: The executive order contributes to the issue’s policy terrain around AI innovation, security, cybersecurity, and frontier systems.
Issue connection: AI security policy; governance terrain.
White House — Fact Sheet
Source type: Government policy summary.
Why it matters: The fact sheet provides a plain-language summary of the executive order’s policy framing.
Issue connection: Policy summary; public-facing context.
06 — Labor, Operator Skill, and Process-Control Work
PwC — 2026 Global AI Jobs Barometer
Link: https://www.pwc.com/gx/en/services/ai/ai-jobs-barometer.html
Source type: Professional-services research report.
Why it matters: PwC’s labor-market analysis informs the issue’s discussion of AI skills, wage signals, productivity, and role changes.
Issue connection: Labor-market terrain; operator skill; inspection and process-control interpretation.
PwC press release — AI reshapes global labor market
Link: https://www.pwc.com/gx/en/news-room/press-releases/2026/pwc-2026-ai-jobs-barometer.html
Source type: Press release.
Why it matters: The release summarizes key findings from the Jobs Barometer for broader public context.
Issue connection: Field spotlight; labor-market signal.
07 — Recursive Self-Improvement Adjacency and Watchlist Context
Cloud Security Alliance — AI Recursive Self-Improvement Security Implications
Source type: Security research / advisory resource.
Why it matters: CSA frames recursive self-improvement as a security-relevant topic, especially around system capability, containment, and risk governance.
Issue connection: Watchlist terrain; recursive-improvement adjacency.
The Economist — How AI got better at building itself
Source type: Journalism / analysis.
Why it matters: The Economist piece contributes to the zeitgeist around AI systems improving AI development processes.
Issue connection: Watchlist context; recursive-improvement public narrative.
08 — Funding, Infrastructure, and Physical-AI Signals
Economic Times — Bezos commits nearly $100M to Flourish
Source type: Funding report.
Why it matters: The article contributes to the issue’s broader terrain scan around capital, physical AI, and brain-inspired AI narratives.
Issue connection: Funding signal; infrastructure and physical-AI terrain.
Dealroom — Bezos leads $500M round in Flourish
Source type: Funding database / market report.
Why it matters: Dealroom adds a second market-facing view of the reported Flourish funding and valuation narrative.
Issue connection: Funding signal; market-terrain watchlist.
09 — Evaluation and Red-Team Tools
promptfoo — LLM red teaming
Link: https://www.promptfoo.dev/docs/red-team/
Source type: Tool documentation.
Why it matters: promptfoo documents red-team testing for LLM applications using simulated adversarial inputs.
Issue connection: Free tools; VSR-02 adversarial testing.
promptfoo GitHub
Link: https://github.com/promptfoo/promptfoo
Source type: Open-source repository.
Why it matters: The repository provides the open-source CLI/library behind promptfoo’s evaluation and red-team workflow.
Issue connection: Free tools; implementation reference.
garak docs
Link: https://docs.garak.ai/garak
Source type: Tool documentation.
Why it matters: garak’s documentation explains its LLM vulnerability-scanning approach.
Issue connection: Free tools; scanner lane.
garak GitHub
Link: https://github.com/NVIDIA/garak
Source type: Open-source repository.
Why it matters: The repository provides source access and project context for garak.
Issue connection: Free tools; security testing.
PyRIT GitHub
Link: https://github.com/Azure/PyRIT
Source type: Open-source repository.
Why it matters: PyRIT is Microsoft’s open-source red-teaming toolkit for generative AI systems.
Issue connection: Free tools; red-team workflow.
10 — Paid Tooling and Product-Scope References
These product links are included to document the tool landscape referenced in VS007. Inclusion does not imply endorsement.
Braintrust
Link: https://www.braintrust.dev/
Source type: Product page.
Issue connection: Evals, logging, prompt management, and AI product iteration.
LangSmith
Link: https://www.langchain.com/langsmith/observability
Source type: Product page.
Issue connection: Tracing, evaluation, observability, and debugging for LLM applications.
Arize Phoenix
Link: https://arize.com/phoenix/
Source type: Product page.
Issue connection: LLM observability and evaluation.
Galileo
Link: https://galileo.ai/
Source type: Product page.
Issue connection: Generative AI evaluation and observability.
Datadog LLM Observability
Link: https://www.datadoghq.com/products/ai/agent-observability/
Source type: Product page.
Issue connection: Agent and LLM observability.
DeepEval / Confident AI
Links: - https://deepeval.com/ - https://www.confident-ai.com/
Source type: Product and framework pages.
Issue connection: LLM evaluation framework and platform.
Helicone
Link: https://www.helicone.ai/
Source type: Product page.
Issue connection: LLM observability and logging.
AgentOps
Link: https://www.agentops.ai/
Source type: Product page.
Issue connection: Agent observability.
Arcade.dev
Link: https://www.arcade.dev/
Source type: Product page.
Issue connection: Agent tool authorization and tool-calling infrastructure.
11 — Work-Context Layers and Capability Surfaces
Microsoft Foundry IQ overview
Link: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq
Source type: Product documentation.
Why it matters: Foundry IQ contributes to the issue’s discussion of work-context layers and how agent systems use enterprise context.
Issue connection: Work-context layer; usefulness and blast-radius tradeoff.
Azure Foundry IQ product page
Link: https://azure.microsoft.com/en-us/products/ai-foundry/iq
Source type: Product page.
Why it matters: The product page gives a public-facing overview of Microsoft’s Foundry IQ direction.
Issue connection: Work-context layer; product terrain.
Microsoft Work IQ overview
Link: https://learn.microsoft.com/en-us/microsoft-365/copilot/extensibility/work-iq/
Source type: Product documentation.
Why it matters: Work IQ provides context for how Microsoft frames workplace knowledge and AI-assisted work.
Issue connection: Work-context layer; enterprise AI surface.
Microsoft Copilot Studio Work IQ overview
Link: https://learn.microsoft.com/en-us/microsoft-copilot-studio/use-work-iq
Source type: Product documentation.
Why it matters: Copilot Studio Work IQ extends the same context-layer terrain into agent and workflow development.
Issue connection: Work-context layer; agent workflow context.
Google DeepMind — Gemini model page
Link: https://deepmind.google/models/gemini/
Source type: Product / model page.
Why it matters: Gemini’s model page is part of the broader field terrain around expanding multimodal and agent-capable AI systems.
Issue connection: Capability-surface expansion.
Google Gemini product page
Link: https://gemini.google/about/
Source type: Product page.
Why it matters: Gemini’s public product page reflects the consumer and operator-facing surface of current model capability.
Issue connection: Capability-surface expansion.
Microsoft AI — Launching seven new MAI models
Link: https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/
Source type: Institutional announcement.
Why it matters: Microsoft’s MAI announcement contributes to the broader field terrain around model specialization and capability surfaces.
Issue connection: Capability-surface expansion.
Microsoft Foundry — New MAI models
Source type: Product announcement.
Why it matters: The Foundry announcement places the MAI models inside Microsoft’s product and deployment ecosystem.
Issue connection: Capability-surface expansion; platform availability.
Anthropic — Claude Design
Link: https://www.anthropic.com/news/claude-design-anthropic-labs
Source type: Institutional announcement.
Why it matters: Claude Design is part of the field’s expanding interface and work-output surface.
Issue connection: Capability-surface expansion; design-facing AI work.
Anthropic support — Get started with Claude Design
Link: https://support.claude.com/en/articles/14604416-get-started-with-claude-design
Source type: Product support documentation.
Why it matters: The support page documents the public user-facing implementation of Claude Design.
Issue connection: Capability-surface expansion; product usage context.
Google Cloud — Eighth-generation TPU for the agentic era
Source type: Institutional infrastructure announcement.
Why it matters: Google Cloud’s TPU announcement contributes to the infrastructure terrain behind larger agentic and model workloads.
Issue connection: AI infrastructure; capability-surface expansion.
Google Cloud — TPU 8t and TPU 8i technical deep dive
Link: https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive
Source type: Technical infrastructure announcement.
Why it matters: The technical deep dive adds infrastructure context for TPU 8t and 8i.
Issue connection: AI infrastructure; technical terrain.
12 — DFEI Reference Resource
The AI Learning Blueprint
Public page: /dispatches/resources/the-ai-learning-blueprint/
PDF: /assets/downloads/resources/DFEI007_The-AI-Learning-Blueprint.pdf
Source type: DFEI companion resource.
Why it matters: The Blueprint provides a role-based operator learning map for Practitioner, Builder, and Architect paths.
Issue connection: Companion learning resource; operator education layer.
13 — DFEI Interpretation Notes
Several issue-level claims combine external source material with DFEI interpretation. In those cases, the source establishes the terrain, while DFEI supplies the operating-frame language.
Examples include:
- “Observability is not safety.”
- “Continuation has become its own control surface.”
- “A repair claim requires verified delta.”
- “A green status must lose to a stop condition.”
- “Labor value shifts toward inspection and process control.”
- “Work-context layers increase usefulness and blast radius.”
These lines are not presented as direct quotations from the linked sources. They are DFEI operating interpretations built from the broader source terrain above.