Research and evidence

Research Ledger

Dated primary-source research, portable findings, attribution, uncertainty, and implementation implications behind The Applied AI Field Guide.

This folder holds dated evidence that supports the implementation library.

Repository source review: The general source set was last rechecked 2026-08-10; the multi-agent topology note was reviewed 2026-08-14; the Bridgewater PAT field report was reviewed 2026-08-27; Uber's production AI operating reports and Warp's software-factory improvement article were reviewed 2026-08-28; the healthcare-claims context, agentic operating-maturity, and FDE interaction-workflow evidence were reviewed 2026-08-30; and the operational-redesign plus workflow-proof and deployment-qualification evidence was reviewed 2026-09-03. Link reachability, first-party publication, or a research preprint does not establish general validity; every implementation pattern still requires local evaluation.

How sources are admitted

Core guidance requires one of:

  • Official protocol specifications, SDK release notes, security advisories, or engineering blogs
  • Reputable research with a reproducible artifact, clear methodology, or directly inspectable benchmark
  • First-party operating reports from identifiable practitioners when the post contains concrete implementation detail
  • Incident reports from the affected organization or a clearly labeled preliminary disclosure

Social, Reddit, YouTube, and news are research channels—not automatic authorities. A post or article enters core guidance only when it is either a qualified first-party field report or points to a primary technical artifact. Otherwise it is recorded as a lead or excluded.

Current research sets

  • Production-agent source ledger — vetted findings reviewed from 2026-02-07 through 2026-08-07, plus explicitly marked canonical sources published earlier and revalidated during that window; includes implementation patterns, anti-patterns, and source-quality notes.
  • AI Engineer production-agent video index — chapter-level practitioner talks with corroboration and claim limits.
  • Operational-redesign and applied-AI practice note — supplied practitioner material, reviewed organizational and company-adoption reports, a directly reviewed change framework, and an inspected open-source curriculum, with portable field, decision-rights, proof, operating, and learning mechanics separated from attribution, forecast, maturity, and production-proof claims.
  • Evidence graphs and change intelligence — architecture/catalog/lineage sources and project implementation leads, scoped to derived maps and impact review rather than graph-driven authority.
  • Reference-solution standards note — primary identity, provisioning, usage-metering, and telemetry anchors for the enterprise-foundation and deployment accelerators, with provider and conformance limits.
  • Business-flow and vertical-solution evidence note — outcome-led solution design, operational applications, healthcare access workflows, financial-investigation evidence, and OT context boundaries used by the business-flow and industry profiles.
  • FDE commercial and professional-practice note — secondary-source leads for pilot graduation, continuation health, delivery-portfolio economics, responsible reuse, and professional boundaries; excludes copied content, commercial adaptation, benchmarks, and pricing doctrine.
  • FDE product-boundaries and capability-transfer note — directly inspectable practitioner evidence for contribution zones, product feedback, shadow-product avoidance, receiving-team capability, and reuse-rights decisions, with organizational claims kept non-normative.
  • Multimodal operations and computer-use boundaries — attributed field evidence for long-running multimodal work, reference-label authority, and browser/desktop action boundaries, with vendor metrics and an insurance vertical profile explicitly deferred.
  • Multi-agent topology selection evidence — controlled evidence for matched-budget single-agent and multi-agent comparison, topology-dependent performance, error containment, and limits on portable thresholds.
  • Healthcare claims context and evaluation — qualified practitioner evidence for data-first delivery, compiled decision context, model-boundary privacy, deterministic adjudication, and change-triggered evaluation, with every engagement metric retained as self-reported.
  • Agentic operating maturity — qualified practitioner evidence for durable execution and lifecycle discipline before orchestration, with the proposed seven-stage ladder and numerical heuristics rejected as defaults.
  • FDE interaction workflow ergonomics — directly inspectable implementation evidence for post-interaction capture, engagement isolation, situation-first entry, reviewed append semantics, and derived readouts, with its method catalog, storage convention, and trust scores rejected as defaults.
  • Bridgewater PAT: governed institutional analysis — bounded first-party field evidence for codified-method context, typed analytical planning, constrained execution, specialist admission, and controlled feedback learning.
  • Uber production AI operating lessons — bounded first-party evidence for causal agent-work cost analysis, direct and downstream model contracts, staged semantic file analysis, and verifiable delegated-agent identity.
  • Warp software factories: governed configuration improvement — bounded vendor evidence for exact configuration baselines, trace scoring, matched configuration comparisons, and protected change proposals, with cloud-only and inherently multi-agent claims rejected as defaults.
  • Workflow proof and deployment qualification — bounded research and practitioner evidence for decision-yielding proofs, participant capacity, baseline acknowledgment, frozen human-AI policy qualification, review burden, cost, and component-sensitive requalification.

Archive

Source fields

Each core-evidence entry records a stable ID, source date or review date, source type, practical finding, and portable pattern. Caveats are included when recency, methodology, attribution, maturity, or generalizability limits how the source should be used. Library pages cite entries by ID so a recommendation can be traced back to its evidence and re-evaluated as platforms change.