Secretary Blinken Tours Argo AI

Abstract

Current proposals in AI governance emphasize mechanisms such as system identification, auditing, and incident reporting. While valuable, these approaches are typically designed as discrete interventions. This creates a structural gap: the absence of a unified framework that links authorization, resource allocation, deployment, and observed outcomes. This paper argues that effective AI accountability requires not isolated traceability points, but an integrated traceability chain. Drawing on established practice in public sector transparency, particularly open contracting and procurement traceability, it outlines a model for linking governance functions across institutional boundaries, and illustrates the cost of fragmented traceability through two documented cases of automated public sector decision making failure. The paper concludes by identifying public sector AI procurement as a tractable starting point for implementing such a chain.

1. Introduction

Recent work in AI governance has proposed a range of mechanisms to improve oversight of advanced systems, including unique identifiers for AI systems, monitoring and visibility tools for deployed agents, third party auditing frameworks, and infrastructure for flaw disclosure. GovAI’s own research programme illustrates this well: its work on system identification, IDs for AI Systems, proposes methods for recognizing and verifying AI systems; its work on agent oversight, Visibility into AI Agents, proposes monitoring tools for systems that act with limited supervision; its work on Frontier AI Auditing proposes third party verification of safety and security claims; and its work on In-House Evaluation Is Not Enough argues for more robust third party flaw disclosure infrastructure for general purpose AI.

Each of these mechanisms addresses a distinct component of the governance problem. However, they are generally developed and implemented independently. As a result, they do not, by default, enable end to end accountability. A basic governance question remains difficult to answer in practice: who authorized a given AI deployment, with what resources, and what were its real world effects?

This paper argues that the difficulty of answering this question is not due to the absence of individual governance tools, but to the absence of a system that connects them.

close up photo of ethernet cables on network switch
Photo by Sergei Starostin on Pexels.com

2. The Limits of Point-Based Traceability

Existing proposals can be understood as forms of point based traceability: system identification enables recognition of a model or deployment; audits evaluate specific claims about safety or performance; monitoring tools provide visibility into system behavior; incident disclosure infrastructure records harms after they occur.

These interventions improve transparency at specific stages of a system’s lifecycle. However, they do not necessarily establish relationships between stages. A system may possess a unique identifier, pass a safety audit, be subject to monitoring, and have a disclosed incident history, while still lacking a coherent account linking the authority that approved deployment, the resources used to build and operate the system, the context in which it was deployed, and the outcomes it produced.

Two documented cases of automated public sector decision making illustrate the cost of this fragmentation.

Australia’s Robodebt scheme (2016 to 2019). The Australian government’s automated welfare debt recovery system compared annualized income data against fortnightly welfare payments and issued automated debt notices where a discrepancy was flagged. The underlying assumption, that recipients had stable and consistent employment, applied to only a small fraction of welfare recipients, which produced systematically inflated debt calculations. A subsequent Royal Commission found the scheme was sustained by a broader institutional failure: the calculation method itself was not transparent to those affected, and the burden of disproving a debt fell on individuals rather than the government having to justify it. Crucially, no single point of accountability existed connecting the ministerial decision to proceed, the department’s calculation methodology, and the outcomes experienced by the roughly 400,000 people wrongly pursued for debts. Each institutional layer could, and did, point elsewhere when the scheme was challenged.

The Dutch childcare benefits scandal and SyRI (2013 to 2021). Dutch tax authorities used a risk scoring algorithm to flag childcare benefit recipients for fraud investigation; a related system, SyRI, aggregated data across agencies to profile welfare fraud risk more broadly. The risk scoring tool used criteria that led to discriminatory targeting based on nationality, ethnicity, and income, and the resulting Childcare Benefit Scandal led to the resignation of the entire Dutch cabinet in 2021 after a parliamentary committee found that fundamental principles of the rule of law had been violated. As in the Robodebt case, the components existed in isolation: a scoring algorithm, a separate flagging registry, and tax administration decisions, without a linked chain that would have let an external reviewer trace a single family’s outcome back to the specific authorization and data inputs that produced it. It took a parliamentary inquiry, not a governance system, to reconstruct that chain after the fact.

Both cases show the same structural failure: the individual components of an accountability system existed, but no shared identifier or institutional linkage connected who authorized the system, what resources and data powered it, and what it did to specific people.

solitary person in empty theater seating
Photo by Fatih Kılıç on Pexels.com

3.Traceability Chains in Public Procurement Governance

Public sector transparency systems have addressed a structurally similar problem in a different domain: linking authority, resources, and outcomes across institutional layers in public spending. The clearest working example is the Open Contracting Data Standard, developed by the Open Contracting Partnership.

The standard is built around the principle that a single contracting process, from planning and tender through award, contract, and implementation, should be traceable end to end through one unique identifier. Intermediate elements of the standard specifically allow publishers to connect procurement data to budget and project records, and to real world outcomes such as locations and delivery milestones. The standard has been adopted by dozens of governments and is endorsed by the G20, with implementations ranging from the EU’s Tender Electronic Daily to Ukraine’s ProZorro and Colombia’s SECOP.

The design features that make the Open Contracting Data Standard work are directly relevant to AI governance:

  • A persistent unique identifier that follows a single process across every institutional stage, rather than each agency maintaining a separate, disconnected record.
  • Publicly accessible data at intermediate stages, not only at the start (tender) or end (delivery), but at each transition in between.
  • Deliberate linkage to project and outcome level data, so spending can be connected to what was actually delivered, not just what was authorized.
  • Multi-stakeholder verification, with civil society and independent researchers able to check disclosed data against real world outcomes rather than relying on self reported agency claims.

This is precisely the structural feature missing from the Robodebt and Dutch childcare cases above: a chain that could be followed by an outside party from authorization to outcome, using a shared reference point.

4. Translating Traceability Chains to AI Governance


The translation from procurement governance to AI governance is direct:

Governance FunctionPublic Procurement (OCDS)AI Governance Equivalent
AuthorityContracting agency and approving officialRegistry of deploying entities and their authorization status
Resource AllocationBudget line and tender awardCompute, funding, and infrastructure provenance
ImplementationContract and delivery milestonesDeployment context, use case, and operational environment
OutcomesVerified delivery and project completionIncident reports, performance data, and third party audit results

The critical requirement is not the existence of each component. GovAI’s existing research programme already addresses several of them individually. What is missing is their integration through a shared identifier and explicit linkage requirements, in the same way the Open Contracting Data Standard links a single tender ID across planning, award, and delivery records.

Concretely, a single system identifier, building on the identification schemes proposed in IDs for AI Systems, could be required to persist across procurement records, model or system registries, deployment disclosures, and incident reporting systems built on the infrastructure proposed in In-House Evaluation Is Not Enough. Such linkage would let an external reviewer move from “which system caused this harm” to “who authorized it, how was it funded, where was it deployed, and what evidence exists of its impact.” This is the exact reconstruction that, in both the Robodebt and Dutch cases, required a multi year public inquiry rather than a queryable record.

5. Implementation Challenges

Implementing traceability chains in AI governance presents non-trivial challenges:

  • Institutional fragmentation. Relevant data is often held by different actors, including private firms, government agencies, and third party auditors, the same fragmentation that delayed accountability in both case studies above.
  • Cross-jurisdictional deployment. AI systems frequently operate across national boundaries, complicating standardization in a way domestic procurement systems do not face to the same degree.
  • Incentive misalignment. Organizations may lack incentives to expose linkages that increase accountability, particularly where marginal risk assessments, as discussed in GovAI’s Assessing Risk Relative to Competitors, already create pressure to under-disclose.
  • Technical standardization. Shared identifiers and interoperable systems require coordination on data standards. The Open Contracting Data Standard itself took over a decade of iterative development to reach its current adoption level, suggesting AI governance should expect a similarly long runway.

These challenges suggest that traceability chains are not a straightforward extension of existing proposals, but require deliberate coordination and governance design.

light painting in close up shot
Photo by Merlin Lightpainting on Pexels.com

6. A Tractable Starting Point: Public-Sector AI Procurement

A practical entry point is public sector AI systems, precisely because governments already maintain, at least in principle, procurement records, budget allocations, and administrative oversight structures compatible with the Open Contracting Data Standard model.

Extending these systems to include AI specific traceability would involve: assigning persistent identifiers to procured AI systems in the same way the standard assigns them to contracts; linking procurement data to deployment registries; requiring disclosure of deployment contexts and use cases at the point of contract award, not only after an incident; and integrating incident reporting tied to the same identifier across the system’s operational life. Independent verification could be supported by civil society organizations, academic researchers, and multi-stakeholder oversight bodies, mirroring the role civil society groups already play in monitoring procurement built on the standard.

Had such a chain existed prior to 2016, an outside reviewer would have been able to trace the Robodebt scheme’s underlying assumptions and calculation methodology directly to the ministerial decision authorizing it, without waiting for a Royal Commission to reconstruct that link after the harm had already occurred.

7. Conclusion

Current AI governance efforts have made significant progress in developing tools for identification, auditing, and monitoring. However, these tools remain largely disconnected. This paper has argued that effective accountability requires a shift from traceability points to traceability chains: integrated systems that link authority, resources, deployment, and outcomes.

Such chains are not hypothetical. The Open Contracting Data Standard demonstrates that this kind of end to end traceability is achievable and durable at scale in public governance, while the Robodebt and Dutch childcare benefits cases demonstrate the human cost of its absence in automated public sector decision making. Adapting the Open Contracting model to AI governance does not require starting from first principles, but rather extending and integrating approaches GovAI has already begun to develop.

As AI systems are increasingly deployed in contexts that directly affect public welfare, including benefits allocation, fraud detection, and health services, the absence of end to end traceability will become a critical governance failure. The central challenge is therefore not whether traceability is possible, but whether governance frameworks will prioritize building systems that connect what is already known.

References

Anderljung, M. et al. IDs for AI Systems. GovAI Research Paper.
https://www.governance.ai/research-paper/ids-for-ai-systems

GovAI. Visibility into AI Agents. GovAI Research Paper.
https://www.governance.ai/research-paper/visibility-into-ai-agents

GovAI. Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies. GovAI Research Paper.
https://www.governance.ai/research-paper/frontier-ai-auditing-toward-rigorous-third-party-assessment-of-safety-and-security-practices-at-leading-ai-companies

GovAI. In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI. GovAI Research Paper.
https://www.governance.ai/research-paper/in-house-evaluation-is-not-enough-towards-robust-third-party-flaw-disclosure-for-general-purpose-ai

GovAI. Assessing Risk Relative to Competitors: An Analysis of Current AI Company Policies. GovAI Research Paper.
https://www.governance.ai/research-paper/assessing-risk-relative-to-competitors-an-analysis-of-current-ai-company-policies

GovAI. Computing Power and the Governance of Artificial Intelligence. GovAI Research Paper.
https://www.governance.ai/research-paper/computing-power-and-the-governance-of-artificial-intelligence

Open Contracting Partnership. Data.
https://www.open-contracting.org/data/

Open Contracting Partnership. Open Contracting Data Standard: Documents and Reports. World Bank.
https://documents1.worldbank.org/curated/en/744551614955316901/pdf/Open-Contracting-Data-Standard.pdf

Open Government Partnership. Anti-Corruption: Open Contracting.
https://www.opengovpartnership.org/open-gov-guide/anti-corruption-open-contracting/

Better Government Lab. Robodebt Case.
https://www.bettergovernmentlab.org/projects/robodebt-case

AIAAIC. Robodebt Welfare Debt Recovery.
https://www.aiaaic.org/aiaaic-repository/ai-algorithmic-and-automation-incidents/robodebt-welfare-debt-recovery

Offreins, N. The Dutch Childcare Benefits Scandal, or “Toeslagenaffaire.” Medium.
https://noortjeoffreins.medium.com/the-dutch-childcare-benefits-scandal-or-toeslagenaffaire-50541da52a0e

OpenGlobalRights. Hollow Rights Victories? Dutch Struggles Against Digital Injustice.
https://www.openglobalrights.org/Dutch-struggles-against-injustice-digital-rights-netherlands/

Yigakpoa L. Ikpae works at the intersection of open data, civic technology, and digital governance, with a focus on transparency, accountability, and inclusive infrastructure.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top