SOURCES

SELECTED SOURCE BACKBONE

VS007 SOURCE BACKBONE

DFEI.007 / VANGUARD SIGNAL 007 — Calibration Drift

DISPATCHES / SOURCES & RESEARCH

This source backbone collects the primary reading paths behind VS007 — Calibration Drift. It is designed as a public source map: a way to see the documentation, research, reporting, security references, product materials, and DFEI interpretation layers that informed the issue.

The backbone is not a claim ledger and not a complete bibliography. It is a curated route through the sources most relevant to the issue’s central question:

What evidence should an AI system produce before it is allowed to continue?


01 — Deployment-Like Evaluation and Runtime Evidence

OpenAI — Deployment Simulation

Link: https://openai.com/index/deployment-simulation/

Source type: Primary institutional documentation.

Why it matters: OpenAI describes evaluation work that moves closer to realistic deployment conditions before release. VS007 uses this as part of the broader shift from static benchmark thinking toward deployment-adjacent testing.

Issue connection: Deployment-like evaluation; benchmark limits; pre-release behavioral simulation.

OpenAI Alignment — Public chat data and real-world misalignment evaluation

Link: https://alignment.openai.com/validating-public-evals/

Source type: Primary institutional research note.

Why it matters: The piece discusses limitations of public evaluations and the value of testing against more realistic conversational evidence.

Issue connection: Public-eval limits; deployment evidence; why benchmark performance is not the same as runtime safety.

OpenAI API Docs — Evaluate agent workflows

Link: https://developers.openai.com/api/docs/guides/agent-evals

Source type: Product documentation.

Why it matters: The documentation shows how agent workflow evaluation can include traces, graders, datasets, and evaluation runs.

Issue connection: Agent evaluation infrastructure; trace-based workflow inspection.


02 — Agent Control and Defense-in-Depth

Google DeepMind — Securing the future of AI agents

Link: https://deepmind.google/blog/securing-the-future-of-ai-agents/

Source type: Primary institutional research roadmap.

Why it matters: DeepMind frames AI control as a defense-in-depth problem for agentic systems, including permissions, monitoring, and behavior verification.

Issue connection: Agent control architecture; safe operation; control before continuation.

Axios — Google DeepMind prepares for rogue AI agents

Link: https://www.axios.com/2026/06/18/google-deepmind-prepares-for-rogue-ai-agents

Source type: News coverage.

Why it matters: Axios provides a secondary reporting lens on DeepMind’s agent-control roadmap and its public-policy significance.

Issue connection: Public terrain; agent-control news signal.


03 — Observability, Tracing, and Telemetry

Microsoft Foundry — Observability in Generative AI

Link: https://learn.microsoft.com/en-us/azure/foundry/concepts/observability

Source type: Product documentation.

Why it matters: Microsoft describes observability for generative AI systems, including traces, tool calls, decisions, and service dependencies.

Issue connection: Runtime telemetry; production traces; visibility into system behavior.

Microsoft Foundry — Agent Service overview

Link: https://learn.microsoft.com/en-us/azure/foundry/agents/overview

Source type: Product documentation.

Why it matters: The Agent Service overview establishes the deployment context for managed AI agents.

Issue connection: Agent workflow deployment; production context.

Microsoft Foundry — Agent tracing concept

Link: https://learn.microsoft.com/en-us/azure/foundry/observability/concepts/trace-agent-concept

Source type: Product documentation.

Why it matters: This documentation describes agent-run tracing across inputs, outputs, tool use, retries, latency, and cost.

Issue connection: Trace evidence; runtime observability; continuation evidence.

Microsoft Foundry — Agent development lifecycle

Link: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/development-lifecycle

Source type: Product documentation.

Why it matters: The lifecycle material connects development, evaluation, tracing, review, and monitoring.

Issue connection: Lifecycle framing for agent evaluation and operational review.

OpenTelemetry — GenAI semantic conventions

Link: https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/

Source type: Technical specification.

Why it matters: OpenTelemetry provides vocabulary for GenAI spans, events, attributes, messages, and instrumentation.

Issue connection: Standardized telemetry; instrumentation language.

OpenTelemetry — Generative AI metrics

Link: https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-metrics/

Source type: Technical specification.

Why it matters: These conventions describe metrics for generative AI systems.

Issue connection: Measurement layer; observability infrastructure.

OpenTelemetry — Semantic conventions for agentic systems discussion

Link: https://github.com/open-telemetry/semantic-conventions-genai/issues/35

Source type: Technical discussion.

Why it matters: The discussion reflects emerging interest in agentic observability conventions.

Issue connection: Agent telemetry watchlist; emerging instrumentation language.

AgentSight — System-Level Observability for AI Agents Using eBPF

Link: https://arxiv.org/abs/2508.02736

Source type: Research paper.

Why it matters: AgentSight explores system-level observability for AI agents, including correlations between semantic actions and system behavior.

Issue connection: Observability gap; agent-system telemetry research.


04 — Security, Red Teaming, and GenAI Risk

OWASP GenAI Security Project

Link: https://genai.owasp.org/

Source type: Security reference.

Why it matters: OWASP’s GenAI Security Project collects security categories and resources for generative AI and agentic applications.

Issue connection: Security reference layer; GenAI application risk.

OWASP Top 10 for LLM Applications

Link: https://owasp.org/www-project-top-10-for-large-language-model-applications/

Source type: Security reference.

Why it matters: The OWASP Top 10 organizes common LLM application risk categories, including risks relevant to excessive agency, prompt injection, and data exposure.

Issue connection: Application-risk framing; agentic workflow risk.

PyRIT paper

Link: https://arxiv.org/abs/2410.02828

Source type: Research paper.

Why it matters: PyRIT presents an automated red-teaming framework for identifying risks in generative AI systems.

Issue connection: Red-team tooling; adversarial testing.

garak paper

Link: https://arxiv.org/abs/2406.11036

Source type: Research paper.

Why it matters: garak provides a structured approach to probing LLM vulnerabilities.

Issue connection: Scanner-based testing; adversarial evaluation.


05 — AI Security Policy and Governance Terrain

White House — Executive Order 14409

Link: https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/

Source type: Government policy document.

Why it matters: The executive order contributes to the issue’s policy terrain around AI innovation, security, cybersecurity, and frontier systems.

Issue connection: AI security policy; governance terrain.

White House — Fact Sheet

Link: https://www.whitehouse.gov/fact-sheets/2026/06/fact-sheet-president-donald-j-trump-promotes-advanced-artificial-intelligence-innovation-and-security/

Source type: Government policy summary.

Why it matters: The fact sheet provides a plain-language summary of the executive order’s policy framing.

Issue connection: Policy summary; public-facing context.


06 — Labor, Operator Skill, and Process-Control Work

PwC — 2026 Global AI Jobs Barometer

Link: https://www.pwc.com/gx/en/services/ai/ai-jobs-barometer.html

Source type: Professional-services research report.

Why it matters: PwC’s labor-market analysis informs the issue’s discussion of AI skills, wage signals, productivity, and role changes.

Issue connection: Labor-market terrain; operator skill; inspection and process-control interpretation.

PwC press release — AI reshapes global labor market

Link: https://www.pwc.com/gx/en/news-room/press-releases/2026/pwc-2026-ai-jobs-barometer.html

Source type: Press release.

Why it matters: The release summarizes key findings from the Jobs Barometer for broader public context.

Issue connection: Field spotlight; labor-market signal.


07 — Recursive Self-Improvement Adjacency and Watchlist Context

Cloud Security Alliance — AI Recursive Self-Improvement Security Implications

Link: https://labs.cloudsecurityalliance.org/research/ai-recursive-self-improvement-security-implications-v1-0-csa/

Source type: Security research / advisory resource.

Why it matters: CSA frames recursive self-improvement as a security-relevant topic, especially around system capability, containment, and risk governance.

Issue connection: Watchlist terrain; recursive-improvement adjacency.

The Economist — How AI got better at building itself

Link: https://www.economist.com/science-and-technology/2026/06/07/how-artificial-intelligence-got-better-at-building-itself

Source type: Journalism / analysis.

Why it matters: The Economist piece contributes to the zeitgeist around AI systems improving AI development processes.

Issue connection: Watchlist context; recursive-improvement public narrative.


08 — Funding, Infrastructure, and Physical-AI Signals

Economic Times — Bezos commits nearly $100M to Flourish

Link: https://startup.economictimes.indiatimes.com/news/funding-deals/bezos-commits-nearly-100m-to-flourish-for-brain-inspired-ai/131546655

Source type: Funding report.

Why it matters: The article contributes to the issue’s broader terrain scan around capital, physical AI, and brain-inspired AI narratives.

Issue connection: Funding signal; infrastructure and physical-AI terrain.

Dealroom — Bezos leads $500M round in Flourish

Link: https://app.dealroom.co/news/note/jeff-bezos-leads-500m-round-in-brain-inspired-ai-startup-flourish-at-2-5b-valuation

Source type: Funding database / market report.

Why it matters: Dealroom adds a second market-facing view of the reported Flourish funding and valuation narrative.

Issue connection: Funding signal; market-terrain watchlist.


09 — Evaluation and Red-Team Tools

promptfoo — LLM red teaming

Link: https://www.promptfoo.dev/docs/red-team/

Source type: Tool documentation.

Why it matters: promptfoo documents red-team testing for LLM applications using simulated adversarial inputs.

Issue connection: Free tools; VSR-02 adversarial testing.

promptfoo GitHub

Link: https://github.com/promptfoo/promptfoo

Source type: Open-source repository.

Why it matters: The repository provides the open-source CLI/library behind promptfoo’s evaluation and red-team workflow.

Issue connection: Free tools; implementation reference.

garak docs

Link: https://docs.garak.ai/garak

Source type: Tool documentation.

Why it matters: garak’s documentation explains its LLM vulnerability-scanning approach.

Issue connection: Free tools; scanner lane.

garak GitHub

Link: https://github.com/NVIDIA/garak

Source type: Open-source repository.

Why it matters: The repository provides source access and project context for garak.

Issue connection: Free tools; security testing.

PyRIT GitHub

Link: https://github.com/Azure/PyRIT

Source type: Open-source repository.

Why it matters: PyRIT is Microsoft’s open-source red-teaming toolkit for generative AI systems.

Issue connection: Free tools; red-team workflow.


10 — Paid Tooling and Product-Scope References

These product links are included to document the tool landscape referenced in VS007. Inclusion does not imply endorsement.

Braintrust

Link: https://www.braintrust.dev/

Source type: Product page.

Issue connection: Evals, logging, prompt management, and AI product iteration.

LangSmith

Link: https://www.langchain.com/langsmith/observability

Source type: Product page.

Issue connection: Tracing, evaluation, observability, and debugging for LLM applications.

Arize Phoenix

Link: https://arize.com/phoenix/

Source type: Product page.

Issue connection: LLM observability and evaluation.

Galileo

Link: https://galileo.ai/

Source type: Product page.

Issue connection: Generative AI evaluation and observability.

Datadog LLM Observability

Link: https://www.datadoghq.com/products/ai/agent-observability/

Source type: Product page.

Issue connection: Agent and LLM observability.

DeepEval / Confident AI

Links: - https://deepeval.com/ - https://www.confident-ai.com/

Source type: Product and framework pages.

Issue connection: LLM evaluation framework and platform.

Helicone

Link: https://www.helicone.ai/

Source type: Product page.

Issue connection: LLM observability and logging.

AgentOps

Link: https://www.agentops.ai/

Source type: Product page.

Issue connection: Agent observability.

Arcade.dev

Link: https://www.arcade.dev/

Source type: Product page.

Issue connection: Agent tool authorization and tool-calling infrastructure.


11 — Work-Context Layers and Capability Surfaces

Microsoft Foundry IQ overview

Link: https://learn.microsoft.com/en-us/azure/foundry/agents/concepts/what-is-foundry-iq

Source type: Product documentation.

Why it matters: Foundry IQ contributes to the issue’s discussion of work-context layers and how agent systems use enterprise context.

Issue connection: Work-context layer; usefulness and blast-radius tradeoff.

Azure Foundry IQ product page

Link: https://azure.microsoft.com/en-us/products/ai-foundry/iq

Source type: Product page.

Why it matters: The product page gives a public-facing overview of Microsoft’s Foundry IQ direction.

Issue connection: Work-context layer; product terrain.

Microsoft Work IQ overview

Link: https://learn.microsoft.com/en-us/microsoft-365/copilot/extensibility/work-iq/

Source type: Product documentation.

Why it matters: Work IQ provides context for how Microsoft frames workplace knowledge and AI-assisted work.

Issue connection: Work-context layer; enterprise AI surface.

Microsoft Copilot Studio Work IQ overview

Link: https://learn.microsoft.com/en-us/microsoft-copilot-studio/use-work-iq

Source type: Product documentation.

Why it matters: Copilot Studio Work IQ extends the same context-layer terrain into agent and workflow development.

Issue connection: Work-context layer; agent workflow context.

Google DeepMind — Gemini model page

Link: https://deepmind.google/models/gemini/

Source type: Product / model page.

Why it matters: Gemini’s model page is part of the broader field terrain around expanding multimodal and agent-capable AI systems.

Issue connection: Capability-surface expansion.

Google Gemini product page

Link: https://gemini.google/about/

Source type: Product page.

Why it matters: Gemini’s public product page reflects the consumer and operator-facing surface of current model capability.

Issue connection: Capability-surface expansion.

Microsoft AI — Launching seven new MAI models

Link: https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/

Source type: Institutional announcement.

Why it matters: Microsoft’s MAI announcement contributes to the broader field terrain around model specialization and capability surfaces.

Issue connection: Capability-surface expansion.

Microsoft Foundry — New MAI models

Link: https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/new-mai-models-in-microsoft-foundry-across-text-image-voice-and-speech/4524632

Source type: Product announcement.

Why it matters: The Foundry announcement places the MAI models inside Microsoft’s product and deployment ecosystem.

Issue connection: Capability-surface expansion; platform availability.

Anthropic — Claude Design

Link: https://www.anthropic.com/news/claude-design-anthropic-labs

Source type: Institutional announcement.

Why it matters: Claude Design is part of the field’s expanding interface and work-output surface.

Issue connection: Capability-surface expansion; design-facing AI work.

Anthropic support — Get started with Claude Design

Link: https://support.claude.com/en/articles/14604416-get-started-with-claude-design

Source type: Product support documentation.

Why it matters: The support page documents the public user-facing implementation of Claude Design.

Issue connection: Capability-surface expansion; product usage context.

Google Cloud — Eighth-generation TPU for the agentic era

Link: https://blog.google/innovation-and-ai/infrastructure-and-cloud/google-cloud/eighth-generation-tpu-agentic-era/

Source type: Institutional infrastructure announcement.

Why it matters: Google Cloud’s TPU announcement contributes to the infrastructure terrain behind larger agentic and model workloads.

Issue connection: AI infrastructure; capability-surface expansion.

Google Cloud — TPU 8t and TPU 8i technical deep dive

Link: https://cloud.google.com/blog/products/compute/tpu-8t-and-tpu-8i-technical-deep-dive

Source type: Technical infrastructure announcement.

Why it matters: The technical deep dive adds infrastructure context for TPU 8t and 8i.

Issue connection: AI infrastructure; technical terrain.


12 — DFEI Reference Resource

The AI Learning Blueprint

Public page: /dispatches/resources/the-ai-learning-blueprint/

PDF: /assets/downloads/resources/DFEI007_The-AI-Learning-Blueprint.pdf

Source type: DFEI companion resource.

Why it matters: The Blueprint provides a role-based operator learning map for Practitioner, Builder, and Architect paths.

Issue connection: Companion learning resource; operator education layer.


13 — DFEI Interpretation Notes

Several issue-level claims combine external source material with DFEI interpretation. In those cases, the source establishes the terrain, while DFEI supplies the operating-frame language.

Examples include:

  • “Observability is not safety.”
  • “Continuation has become its own control surface.”
  • “A repair claim requires verified delta.”
  • “A green status must lose to a stop condition.”
  • “Labor value shifts toward inspection and process control.”
  • “Work-context layers increase usefulness and blast radius.”

These lines are not presented as direct quotations from the linked sources. They are DFEI operating interpretations built from the broader source terrain above.