diff --git a/staging/batches/BATCH-2026-009/claim-lineage.csv b/staging/batches/BATCH-2026-009/claim-lineage.csv new file mode 100644 index 0000000..1175d94 --- /dev/null +++ b/staging/batches/BATCH-2026-009/claim-lineage.csv @@ -0,0 +1,11 @@ +"claim_id","claim","claim_type","lineage","calculation","use_status" +"GAP-001","No production dossier can be drafted for this topic.","editorial threshold conclusion","METHODOLOGY.md; config/publication-thresholds.yml; canonical data inventory","0 production sources versus 30 required","research-gap report only" +"GAP-002","The canonical topic corpus has zero sources, statements, propositions, people, books, and dossiers.","inventory calculation","data/sources; data/statements; data/propositions; data/people; data/dossiers; book directories","Count non-README production records","research-gap report only" +"GAP-003","Eight intellectual sources currently have unpublished statements.","staging gap calculation","BATCH-2026-005 source-proposals.json; BATCH-2026-006 statements.json","9 source records minus 1 duplicate rendition with no created statements","not dossier evidence" +"GAP-004","The unpublished statement layer contains 46 records with human review pending.","staging gap calculation","BATCH-2026-006 manifest.yml; statements.json","Count statements and inspect human_review_status","not dossier evidence" +"GAP-005","Three unpublished primary empirical works are synthetic or sandboxed.","evidence-character calculation","Statements 022–038 and source records 006–008","Distinct primary empirical source IDs","not dossier evidence" +"GAP-006","The staging corpus does not support role-specific comparisons.","coverage conclusion","BATCH-2026-006 statements.json","All 46 share relevance tags; direct role-specific position records = 0","research-gap report only" +"GAP-007","The staging corpus does not support regional comparisons.","coverage conclusion","BATCH-2026-005 source-proposals.json; BATCH-2026-006 statements.json","No region-comparative or region-specific empirical population","research-gap report only" +"GAP-008","No material disagreement is production-ready.","coverage conclusion","data/stances inventory; BATCH-2026-006 statement types","0 production stances and 0 counterargument statements in staging","research-gap report only" +"GAP-009","OFF-owned source share is undefined in production and zero in the current staging source set.","ownership calculation","Canonical source inventory; BATCH-2026-005 source-proposals.json","0/0 production; 0/8 staging intellectual sources","not dossier evidence" +"GAP-010","Existing OFF web content overlaps the topic but has not been ingested into the index.","overlap review","OFF CISO AI Leverage Report and related official pages, checked 2026-08-15","Relevant web results minus repository source records","duplication and next-batch planning only" diff --git a/staging/batches/BATCH-2026-009/content-quality-review.md b/staging/batches/BATCH-2026-009/content-quality-review.md new file mode 100644 index 0000000..6f4a664 --- /dev/null +++ b/staging/batches/BATCH-2026-009/content-quality-review.md @@ -0,0 +1,23 @@ +# Content quality review + +## Result + +Pass for a research-gap report; public dossier intentionally withheld. + +## Checks + +- Generic opening removed: the report begins with the threshold decision and denominator. +- Unsupported urgency removed: no claim of unprecedented change, inevitability, market momentum, or universal adoption appears. +- Artificial symmetry avoided: missing evidence categories reflect actual corpus gaps rather than balanced-for-style sections. +- Source summaries not stacked: staging works are grouped by evidence character and discussed only as gaps. +- Consensus language avoided: the report uses “recurring normative alignment within a selected staging corpus” and explicitly denies universal inference. +- Weakly supported ideas are framed as testable questions, not mocked or presented as facts. +- Role claims calibrated: shared relevance tags are not interpreted as role beliefs. +- Regional claims calibrated: synthetic study affiliation is not treated as market or study geography. +- OFF overlap disclosed: official pages are described as unverified external-to-repository content and not counted as evidence. +- Repetition controlled: threshold numbers appear only where required for decision, corpus table, and next-step calculation. +- No public dossier headings or padded executive implications were drafted after threshold failure. + +## Human review requested + +Confirm the chosen research question, the production inventory, the classification of staging material, the OFF overlap disposition, and the proposed evidence-acquisition sequence. diff --git a/staging/batches/BATCH-2026-009/dossier-protocol.yml b/staging/batches/BATCH-2026-009/dossier-protocol.yml new file mode 100644 index 0000000..77e3767 --- /dev/null +++ b/staging/batches/BATCH-2026-009/dossier-protocol.yml @@ -0,0 +1,41 @@ +protocol_id: dossier-protocol-BATCH-2026-009-001 +batch_id: BATCH-2026-009 +prompt_id: OEII-TOPIC-DOSSIER +prompt_version: "2.0" +status: pre_registered_threshold_failed +topic: Governed identities for AI agents +research_question: How should organizations govern identity, delegated authority, security evaluation, and accountability for AI agents? +scope: Cross-sector organizational controls for identifying AI agents, tracing represented principals, bounding delegated authority, managing credential lifecycles, evaluating security, and allocating accountability. +time_period: 2020-08-15 through 2026-08-15, with pre-2020 foundational identity, authorization, and zero-trust material eligible only when directly governing current agent questions. +included_propositions: [] +included_proposition_rule: Canonical production proposition, published workflow state, complete statement lineage, and appropriate human review. +excluded_adjacent_questions: + - General generative-AI governance without agent identity, authority, or accountability relevance + - Human IAM practices with no demonstrated applicability to autonomous or delegated workloads + - General AI adoption, ROI, labor effects, or model capability rankings + - Market-size forecasts and vendor rankings + - Production prevalence inferred from synthetic security benchmarks +evidence_hierarchy: + - Independently verified production empirical evidence with reproducible methods and statement lineage + - Adopted standards, public policy, and verified legal or regulatory materials + - Independently checked operator implementation evidence + - Normative frameworks and expert recommendations with explicit attribution + - Company-reported outcomes, kept attributed and independently checked where possible + - OFF-owned evidence with ownership, selection, sponsorship, and sample disclosures +comparison_method: Compare evidence character, scope, technology, autonomy, role, geography, industry, time horizon, and independence; source frequency is not truth. +role_analysis_method: Require direct, independently checked role-specific sources; shared executive-relevance tags do not establish what a role believes. +regional_analysis_method: Separate speaker, organization, market, study, and regulatory geography; require at least two people, two organizations, and region-specific sources for a regional briefing. +trend_method: Use frozen releases, stable eligibility rules, disclosed denominators, and no momentum claim without a configured calculation. +expected_limitations: + - Production corpus is empty for the topic + - Pending staging evidence is English-only and weak on regional and role-specific comparison + - Empirical security evidence is synthetic or sandboxed + - Normative sources do not report comparative implementation outcomes +criteria_for_editorial_conclusions: + - Threshold must pass before drafting + - Every substantive claim must resolve to production statement, proposition, source, or named calculation + - Evidence character and uncertainty must be stated + - Material counterpositions and concentration must be disclosed + - Named human approval is required +threshold_disposition: Do not draft content/topic-dossiers or data/dossiers; create a staging research-gap package and recommended next batches. +human_review_status: pending diff --git a/staging/batches/BATCH-2026-009/evidence-matrix.csv b/staging/batches/BATCH-2026-009/evidence-matrix.csv new file mode 100644 index 0000000..ed8e3eb --- /dev/null +++ b/staging/batches/BATCH-2026-009/evidence-matrix.csv @@ -0,0 +1,7 @@ +"matrix_id","layer","record_ids","count","evidence_character","production_eligible","permitted_use","principal_limit","next_action" +"E-001","Canonical sources","none","0","none","no","Threshold calculation only","No production evidence exists","Human-review and promote eligible records; add missing sources" +"E-002","Unpublished normative and standards sources","source-BATCH-2026-005-001; source-BATCH-2026-005-002; source-BATCH-2026-005-003; source-BATCH-2026-005-004; source-BATCH-2026-005-009","5","RFI synthesis, public guidance, draft standard, standards report, vendor governance report","no","Research-gap planning only","Normative or interpretive; no comparative implementation outcomes; human review pending","Run named human review and add operator implementation evidence" +"E-003","Unpublished primary synthetic empirical sources","source-BATCH-2026-005-006; source-BATCH-2026-005-007; source-BATCH-2026-005-008","3","Synthetic prompt-injection, agent-security, and sandboxed vulnerability experiments","no","Research-gap planning only","Does not estimate production prevalence; human review pending","Add independent production incident and deployment studies" +"E-004","Unpublished statements","statement-BATCH-2026-006-001 through statement-BATCH-2026-006-046","46","Mixed: empirical, recommendation, warning, definition, framework, interpretation, methodology, policy","no","Coverage and gap planning only","All human-review statuses pending","Complete named statement and attribution review" +"E-005","Candidate-only proposition ideas","proposition-candidate-BATCH-2026-006-001 through proposition-candidate-BATCH-2026-006-009","9","Unnormalized candidate links","no","Gap planning only","Not canonical propositions and not human approved","Complete proposition and stance review before promotion" +"E-006","Potential OFF overlap","CISO AI Leverage Report; Executive AI Leverage Report; CISO Executive Forum pages","3 content families","OFF-owned operator research and community content","no","Duplication review and ingestion planning only","Outside repository corpus; self-selected samples and commercial relationships require disclosure","Create a separate OFF-content verification and rights batch" diff --git a/staging/batches/BATCH-2026-009/included-records.json b/staging/batches/BATCH-2026-009/included-records.json new file mode 100644 index 0000000..13a4feb --- /dev/null +++ b/staging/batches/BATCH-2026-009/included-records.json @@ -0,0 +1,24 @@ +{ + "batch_id": "BATCH-2026-009", + "production_source_ids": [], + "production_statement_ids": [], + "production_proposition_ids": [], + "production_person_ids": [], + "production_book_ids": [], + "production_dossier_ids": [], + "discovery_protocol_ids": [ + "protocol-BATCH-2026-001" + ], + "staging_records_used_as_evidence": [], + "staging_records_referenced_as_unpublished_research_gaps": { + "source_batch": "BATCH-2026-005", + "statement_batch": "BATCH-2026-006", + "person_batch": "BATCH-2026-007" + }, + "external_sources_used_as_evidence": [], + "off_pages_reviewed_for_overlap_only": [ + "https://openfutureforum.com/research/ciso-ai-leverage-report", + "https://openfutureforum.com/research/executive-ai-leverage-report", + "https://openfutureforum.com/ciso-executive-forum" + ] +} diff --git a/staging/batches/BATCH-2026-009/limitations.md b/staging/batches/BATCH-2026-009/limitations.md new file mode 100644 index 0000000..9333a57 --- /dev/null +++ b/staging/batches/BATCH-2026-009/limitations.md @@ -0,0 +1,15 @@ +# Limitations + +- The production corpus is empty; this package cannot answer the research question. +- Staging counts are reported only to plan research and may change after named human review, deduplication, or rejection. +- The staging source set is English-only and contains no region-comparative evidence. +- Shared executive-relevance tags do not constitute role-specific testimony or measurement. +- Synthetic and sandboxed security studies do not estimate production prevalence or control effectiveness. +- Normative standards and governance documents do not establish implementation outcomes. +- The source set contains vendor involvement; source count is not treated as truth. +- OFF web content was checked through public search on 2026-08-15 for overlap only and was not independently verified or ingested. +- No historical release exists for trend analysis. +- No book record is available for a foundational-work comparison. +- No formal counterargument statement or production stance record is available for debate mapping. +- The next-batch targets are editorial recommendations, not claims that the identified sources or respondents will be obtainable. +- All records remain machine-produced and human-review pending. diff --git a/staging/batches/BATCH-2026-009/manifest.yml b/staging/batches/BATCH-2026-009/manifest.yml new file mode 100644 index 0000000..42697ed --- /dev/null +++ b/staging/batches/BATCH-2026-009/manifest.yml @@ -0,0 +1,61 @@ +batch_id: BATCH-2026-009 +prompt_id: OEII-TOPIC-DOSSIER +prompt_version: "2.0" +topic: Governed identities for AI agents +topic_slug: governed-agent-identities +research_question: How should organizations govern identity, delegated authority, security evaluation, and accountability for AI agents? +time_period: 2020-08-15 through 2026-08-15, with pre-2020 foundations only when directly applicable +branch: analysis/dossier-governed-agent-identities-BATCH-2026-009 +base_branch: research/people-BATCH-2026-007 +execution_date: 2026-08-15 +execution_completed_at: 2026-08-15T16:30:00-07:00 +agent_or_researcher: OpenAI Codex; machine-assisted production-threshold and research-gap review; named human review pending +model_disclosure: AI-assisted canonical inventory, threshold calculation, gap synthesis, matrix drafting, OFF overlap search, and validation; no human or publication approval occurred. +threshold_result: FAIL +disposition: research_gap_report_only +production_sources: 0 +production_people: 0 +production_organizations: 0 +production_source_types: 0 +production_roles: 0 +production_regions: 0 +production_languages: 0 +production_books: 0 +production_empirical_sources: 0 +production_propositions: 0 +production_challenging_positions: 0 +production_statements: 0 +public_dossier_created: false +canonical_dossier_record_created: false +human_review_status: pending +output_files: + - manifest.yml + - threshold-check.json + - dossier-protocol.yml + - research-gap-report.md + - evidence-matrix.csv + - role-matrix.csv + - regional-matrix.csv + - claim-lineage.csv + - included-records.json + - unresolved-questions.json + - original-research-opportunities.json + - revision-record.json + - off-content-overlap-review.md + - content-quality-review.md + - limitations.md + - validation-results.md +intentionally_omitted_outputs: + - content/topic-dossiers/governed-identities-for-ai-agents.md + - data/dossiers/dossier-governed-identities-for-ai-agents.json +omission_reason: Full-topic dossier threshold failed; staging records may not be used as dossier evidence. +validation_required: + - publication threshold + - statement and proposition lineage + - unsupported claims and evidence character + - source concentration and OFF ownership + - role and regional representation + - consensus language and near duplication + - existing OFF content overlap and AI-slop review + - structured data and internal links + - public build and staging isolation diff --git a/staging/batches/BATCH-2026-009/off-content-overlap-review.md b/staging/batches/BATCH-2026-009/off-content-overlap-review.md new file mode 100644 index 0000000..c264201 --- /dev/null +++ b/staging/batches/BATCH-2026-009/off-content-overlap-review.md @@ -0,0 +1,17 @@ +# OFF content overlap review + +Official Open Future Forum web pages were checked on 2026-08-15 for duplication and future-ingestion planning only. They were not added to the evidence corpus. + +## Relevant content found + +- The [CISO AI Leverage Report](https://openfutureforum.com/research/ciso-ai-leverage-report) discusses agent identity and access, security ownership, and a directional governance-budget finding. +- The [Executive AI Leverage Report](https://openfutureforum.com/research/executive-ai-leverage-report) includes a directional security-room finding about agent access. +- The [CISO Executive Forum](https://openfutureforum.com/ciso-executive-forum) and related community pages describe relevant practitioner communities and commercial participants. + +## Disposition + +These pages are not repository evidence and remain outside the verified source corpus. They must not fill the dossier threshold until a dedicated ingestion batch verifies methods, exact statement locators, respondent bases, role classification, dates, ownership, sponsorship, selection effects, and any advisory relationships. OFF ownership would need prominent disclosure. The future dossier should link selectively rather than reproduce these reports. + +## Duplication risk + +A future dossier should not repeat the “governance budget gap” article as its framing. The index contribution should instead compare that operator finding with independent implementation, standards, legal, technical, role-specific, and regional evidence. diff --git a/staging/batches/BATCH-2026-009/original-research-opportunities.json b/staging/batches/BATCH-2026-009/original-research-opportunities.json new file mode 100644 index 0000000..8004894 --- /dev/null +++ b/staging/batches/BATCH-2026-009/original-research-opportunities.json @@ -0,0 +1,49 @@ +{ + "batch_id": "BATCH-2026-009", + "opportunities": [ + { + "opportunity_id": "OFF-OPP-001", + "title": "Agent identity governance ownership study", + "question": "Which executive role owns agent inventory, credentials, authorization policy, incident response, and budget?", + "method": "Role-tagged survey plus follow-up interviews with CISOs, CIOs, CTOs, GCs, CFOs, and board directors.", + "minimum_design": "State respondent counts by role; separate operators from vendors and advisors; publish instrument and denominators.", + "independence_controls": "Disclose OFF recruitment, event participation, sponsorship, and advisory relationships; sponsors may not shape questions or analysis.", + "status": "proposed_for_human_review", + "publication_status": null, + "human_review_status": "pending" + }, + { + "opportunity_id": "OFF-OPP-002", + "title": "Agent authorization implementation registry", + "question": "Which identity and delegated-authorization patterns are deployed, at what scale, and with what failure modes?", + "method": "Structured operator case series with architecture, scale, incident, revocation, audit, and interoperability fields.", + "minimum_design": "At least 12 deployments across 6 organizations and 3 industries; include negative and abandoned implementations.", + "independence_controls": "No pay-to-include entries; vendor claims remain attributed until independently checked.", + "status": "proposed_for_human_review", + "publication_status": null, + "human_review_status": "pending" + }, + { + "opportunity_id": "OFF-OPP-003", + "title": "Human approval and monitoring burden study", + "question": "Where do human approval, hard limits, automated monitoring, and shutdown controls work or fail?", + "method": "Prospective workflow study measuring approval frequency, override rates, response time, false positives, missed events, and user burden.", + "minimum_design": "Pre-registered outcomes and task-risk strata; preserve unsuccessful results.", + "independence_controls": "Independent statistical review and privacy assessment.", + "status": "proposed_for_human_review", + "publication_status": null, + "human_review_status": "pending" + }, + { + "opportunity_id": "OFF-OPP-004", + "title": "Benchmark-to-production validity study", + "question": "Do agent-security benchmark results predict production incidents or control effectiveness?", + "method": "Map benchmark tasks and defenses to anonymized production incidents and red-team exercises.", + "minimum_design": "Multiple models, agent architectures, industries, and time-stamped versions; report denominators and uncertainty.", + "independence_controls": "Separate benchmark authors, vendors, and evaluating operators where possible.", + "status": "proposed_for_human_review", + "publication_status": null, + "human_review_status": "pending" + } + ] +} diff --git a/staging/batches/BATCH-2026-009/regional-matrix.csv b/staging/batches/BATCH-2026-009/regional-matrix.csv new file mode 100644 index 0000000..2347dce --- /dev/null +++ b/staging/batches/BATCH-2026-009/regional-matrix.csv @@ -0,0 +1,10 @@ +"region_or_context","production_sources","unpublished_context_sources","speaker_geography","organization_geography","market_or_regulatory_geography","study_geography","languages","assessment","next_evidence" +"United States","0","2","Not established for institutional records","NIST: United States","United States with global applicability","RFI respondents not sufficiently described","English","Insufficient for a regional claim","US operator implementations plus applicable federal and state legal analysis" +"Global standards","0","2","Mixed or not reported","IETF and OpenID global standards contexts","Global standards context","No deployment population","English","Normative context, not geographic representation","Adoption and interoperability evidence from organizations in multiple regions" +"Global vendor policy","0","1","Authors listed; affiliations incompletely reported","OpenAI publisher","Global policy context","No empirical study geography","English","Vendor-authored normative perspective","Independent legal, civil-society, buyer, and regulator positions" +"Synthetic/no human geography","0","3","Academic author teams","US, European, and Asian institutions appear in affiliations","No market represented","Curated synthetic or sandbox environments","English","Cannot support a regional comparison","Production studies with explicit organization, market, and incident geography" +"European Union and EEA","0","0","none","none","none","none","none","Absent","EU AI Act, cybersecurity, data-protection, and operator implementation sources" +"United Kingdom","0","0","none","none","none","none","none","Absent","UK regulatory and operator sources" +"Asia-Pacific","0","0","none","none","none","none","none","Absent","Singapore, Japan, Australia, India, and other region-specific sources" +"Latin America, Middle East, and Africa","0","0","none","none","none","none","none","Absent","Region-specific regulatory, operator, and civil-society sources" +"Non-English evidence","0","0","none","none","none","none","none","Absent","At least one independently verified non-English source in each selected regional batch" diff --git a/staging/batches/BATCH-2026-009/research-gap-report.md b/staging/batches/BATCH-2026-009/research-gap-report.md new file mode 100644 index 0000000..2cebc0d --- /dev/null +++ b/staging/batches/BATCH-2026-009/research-gap-report.md @@ -0,0 +1,95 @@ +# Research gap report: governed identities for AI agents + +Status: **Full-topic dossier threshold not met; no dossier drafted.** + +## Decision + +The production corpus contains no verified source, statement, proposition, person, book, or prior dossier record for this topic. The configured full-topic dossier threshold requires 30 verified sources, 15 people, five organizations, three source types, independent verification, and named human approval. Every requirement fails. Creating a 4,000–8,000-word dossier would therefore substitute unpublished research for production evidence and contradict the repository methodology. + +This batch creates a research-gap package only. It does not create a file in `content/topic-dossiers/` or a record in `data/dossiers/`. + +## Research question and scope + +The intended question is: **How should organizations govern identity, delegated authority, security evaluation, and accountability for AI agents?** + +The eligible period follows the existing discovery protocol: 15 August 2020 through 15 August 2026, with older identity, authorization, workload, and zero-trust foundations eligible only when they directly govern current agent questions. General generative-AI governance, market forecasts, vendor rankings, and human IAM with no demonstrated agent relevance remain outside scope. + +## Production corpus + +| Measure | Production count | Dossier requirement | +|---|---:|---:| +| Verified sources | 0 | 30 | +| People | 0 | 15 | +| Organizations | 0 | 5 | +| Source types | 0 | 3 | +| Statements | 0 | Required for claim lineage | +| Propositions | 0 | Required for normalized synthesis | +| Books | 0 | No numeric threshold, but none available | +| Empirical sources | 0 | No numeric threshold, but required for evidence sections | +| Challenging positions | 0 | No genuine disagreement can be mapped | + +Ownership, vendor, and consulting shares are undefined for an empty production denominator. No percentage is reported as zero when the denominator is zero. + +## What exists outside production + +The repository’s unpublished research shows where the next work should start, but staging records are not evidence for a public dossier. The staging layer contains nine source records representing eight intellectual works with statements; one record is a duplicate rendition. Those works span academic papers, a public-policy document, research reports, and a working paper. They have 68 distinct named authors, while 15 people have enriched staging records. Forty-six statements remain human-review pending. Nine earlier proposition ideas are explicitly candidate-only, and there are no canonical propositions. + +Four unpublished sources contain empirical findings. Three are primary security studies in synthetic or sandboxed environments; the fourth is a NIST synthesis of a self-selected request-for-information corpus. The synthetic studies can potentially establish bounded capabilities, attack outcomes, and evaluation tradeoffs. They cannot estimate production incident prevalence or show that a control improves outcomes in deployed organizations. The normative standards and governance reports propose useful control models but do not compare implementations. + +The staging source set is English-only. Its normative contexts are United States or global; its primary experiments have no human geography. All 46 statements carry the same five executive-relevance tags, but those tags are not direct evidence of what CISOs, CIOs, CTOs, general counsel, or board directors think. CFO and CEO perspectives are absent from the tags. No role-comparative conclusion is supportable. + +## Strongest potential evidence—and why it is not yet usable + +The strongest potential empirical material is the three-study synthetic security cluster: AgentDojo separates benign utility, utility under attack, and attacker success; ASB compares attack and defense conditions across agent scenarios; HPTSA reports a bounded multi-agent vulnerability-exploitation capability. Their exact staging statements are `statement-BATCH-2026-006-022` through `-038`. These records are promising because they retain methods, scopes, and limitations. They remain unpublished and human-review pending, and their external validity is unresolved. + +The strongest potential governance material connects distinct agent identity, principal traceability, delegated authority, lifecycle control, auditability, monitoring, shutdown, and human legal-entity accountability. The relevant unpublished statements are distributed across NIST, an IETF Internet-Draft, an OpenID Foundation report, and an OpenAI-authored governance report. This is recurring normative alignment within a selected staging corpus, not proof of effectiveness or a population-wide conclusion. + +## Exact evidence missing + +1. **Production promotion.** The eight intellectual works with statements need named source, statement, attribution, rights, and proposition review before they can count. Even if all eight pass, the dossier would still need at least 22 additional verified sources. +2. **Implemented controls.** The corpus needs independent operator cases showing how identity granularity, credential issuance, rotation, revocation, policy enforcement, audit, monitoring, and shutdown behave in production—including failures and abandoned designs. +3. **Outcome evidence.** It needs measured unauthorized-action rates, attack success, task utility, incident response, false-positive burden, recovery time, control cost, and interoperability outcomes. +4. **Counterpositions.** It needs comparable arguments about privacy, anonymity, surveillance, centralization, usability, innovation cost, principal-linkage limits, and the risks of over-broad monitoring. No current statement is coded as a counterargument. +5. **Role evidence.** It needs direct CISO, CIO, CTO, general counsel, CFO, CEO, and board evidence rather than shared relevance tags. +6. **Regional evidence.** It needs explicit regulatory and operating evidence from the European Union, United Kingdom, Asia-Pacific, and underrepresented regions, with speaker, organization, market, study, and regulatory geography separated. +7. **Industry evidence.** It needs implementations from financial services, healthcare, critical infrastructure, public sector, professional services, and technology/cloud environments. +8. **Foundational works.** It needs verified books or foundational standards that clarify how workload identity, zero trust, capability security, delegation, and accountability apply—or fail to apply—to agents. +9. **Longitudinal evidence.** It needs frozen releases before any change-over-time claim can be made. +10. **Independent review and human approval.** These are configured requirements, not editorial options. + +## Weakly supported ideas to test + +- **“Every agent needs a unique identity.”** The staging corpus repeats versions of this recommendation, but it does not compare per-instance, per-session, per-class, or privacy-preserving designs. Interoperability and incident evidence would test it. +- **“Short-lived credentials solve agent credential risk.”** A draft standard recommends short validity, but no comparative production outcome is available. Rotation, revocation latency, availability, and task-failure measurements would test it. +- **“Human approval keeps high-stakes agents safe.”** One governance source also warns that approval for every action can create consent fatigue. Risk-stratified user studies and audit data would test when approval is meaningful. +- **“Automated monitors can supervise agents at machine speed.”** The recommendation is plausible but carries unmeasured reliability, privacy, cost, and control risks. Independent evaluations with missed-event and false-positive rates would test it. +- **“Synthetic benchmark leadership predicts safer deployment.”** Current studies do not establish that relationship. Paired benchmark and production-incident data would test it. + +These ideas are not dismissed; they are separated from established evidence until the appropriate tests exist. + +## Material disagreement currently missing + +There is no production-ready disagreement. Apparent tensions in staging usually concern different scopes: approval for every high-volume action versus approval for high-consequence action; bounded sandbox capability versus production prevalence; and different models producing different benchmark results. A future debate map needs at least 12 relevant statements, two materially distinct positions, comparable scope, and documented limitations. + +## OFF content relationship + +Official OFF pages already discuss agent identity, access, governance ownership, security budgets, and CISO communities. They were reviewed only for overlap. They are not repository evidence and should not be copied into a dossier. A dedicated batch should verify methods, exact statement locators, sample composition, role tags, ownership, sponsorship, and advisory disclosures before any OFF material enters the index. The eventual dossier should add independent comparison rather than repackage the existing CISO AI Leverage Report. + +## Recommended next batches + +1. **Human-review and promotion batch:** review the eight intellectual works, 46 statements, author relationships, and proposition candidates; promote only records that satisfy schema, attribution, rights, independence, and review requirements. +2. **Operator implementation batch:** target at least eight independent production sources across six organizations and three industries, including negative or abandoned implementations. +3. **Counterposition and civil-liberties batch:** target at least four sources on privacy, anonymity, surveillance, usability, over-control, and competing accountability models. +4. **Regional legal and standards batch:** target at least six sources spanning the EU/EEA, UK, and Asia-Pacific, including non-English primary material where feasible. +5. **Industry case batch:** target at least six sources across financial services, healthcare, critical infrastructure, public sector, professional services, and technology/cloud. +6. **Foundational works batch:** verify at least two books or foundational standards with direct agent applicability and stable edition/version locators. +7. **OFF evidence-ingestion batch:** separately verify relevant OFF reports and event research, with ownership and commercial disclosures and no threshold credit until review is complete. +8. **Threshold recheck:** rerun the dossier calculation only after at least 30 sources are production-eligible, proposition lineage is canonical, independence has been checked, and a named human reviewer is available. + +## Recommended OFF original research + +The highest-value follow-up is a role-tagged study of who owns agent identity governance across security, technology, legal, finance, executive, and board functions. It should pair decision rights with actual controls, budgets, incidents, and metrics. A second priority is an implementation registry that records architecture, scale, identity granularity, delegation, revocation, audit, monitoring, and failure modes. Both designs must separate operators from vendors and advisors and disclose OFF recruitment, sponsorship, event participation, and advisory relationships. + +## Human-review requirement + +No record in this package is human approved. A named reviewer must confirm the threshold calculation, staging-versus-production boundary, gap priorities, OFF overlap disclosure, and next-batch design. Only after the production threshold passes should an editor create the public Markdown dossier and canonical dossier JSON record. diff --git a/staging/batches/BATCH-2026-009/revision-record.json b/staging/batches/BATCH-2026-009/revision-record.json new file mode 100644 index 0000000..bb494af --- /dev/null +++ b/staging/batches/BATCH-2026-009/revision-record.json @@ -0,0 +1,13 @@ +{ + "batch_id": "BATCH-2026-009", + "record_type": "research_gap_package", + "revisions": [ + { + "revision_id": "revision-BATCH-2026-009-001", + "changed_at": "2026-08-15T16:30:00-07:00", + "changed_by": "OpenAI Codex", + "summary": "Created production-only threshold calculation and research-gap package; withheld public dossier and canonical dossier record because every configured dossier threshold failed.", + "human_review_status": "pending" + } + ] +} diff --git a/staging/batches/BATCH-2026-009/role-matrix.csv b/staging/batches/BATCH-2026-009/role-matrix.csv new file mode 100644 index 0000000..61f3b38 --- /dev/null +++ b/staging/batches/BATCH-2026-009/role-matrix.csv @@ -0,0 +1,8 @@ +"role_id","production_sources","unpublished_shared_relevance_tags","direct_role_specific_positions","threshold_interpretation","missing_evidence","recommended_minimum_next_sample" +"role-ciso","0","46","0","No claim about CISO views is permitted","Direct CISO problem framing, control ownership, budget, implementation, metrics, and board reporting","3 independently checked CISO/operator sources across at least 2 organizations" +"role-cio","0","46","0","No claim about CIO views is permitted","Identity platform ownership, integration, lifecycle operations, service reliability, and procurement","3 independently checked CIO/operator sources across at least 2 organizations" +"role-cto","0","46","0","No claim about CTO views is permitted","Architecture, developer controls, protocol interoperability, deployment constraints, and shutdown design","3 independently checked CTO/operator sources across at least 2 organizations" +"role-general-counsel","0","46","0","No claim about general-counsel views is permitted","Liability, delegation, privacy, evidence retention, regulatory obligations, and contracting","3 independently checked GC/legal sources across at least 2 organizations or jurisdictions" +"role-board-director","0","46","0","No claim about board views is permitted","Risk appetite, accountability, assurance, audit evidence, materiality, and oversight cadence","3 independently checked board or board-governance sources across at least 2 organizations" +"role-cfo","0","0","0","Role absent from supplied relevance tags","Budget ownership, control cost, insurance, loss measurement, and investment gates","3 independently checked CFO/finance sources across at least 2 organizations" +"role-ceo","0","0","0","Role absent from supplied relevance tags","Accountable executive ownership, deployment pace, risk appetite, and cross-functional decision rights","3 independently checked CEO/operator sources across at least 2 organizations" diff --git a/staging/batches/BATCH-2026-009/threshold-check.json b/staging/batches/BATCH-2026-009/threshold-check.json new file mode 100644 index 0000000..5f0860f --- /dev/null +++ b/staging/batches/BATCH-2026-009/threshold-check.json @@ -0,0 +1,121 @@ +{ + "batch_id": "BATCH-2026-009", + "prompt_id": "OEII-TOPIC-DOSSIER", + "prompt_version": "2.0", + "topic": "Governed identities for AI agents", + "research_question": "How should organizations govern identity, delegated authority, security evaluation, and accountability for AI agents?", + "time_period": "2020-08-15 through 2026-08-15, with pre-2020 foundations only when directly applicable", + "calculation_date": "2026-08-15", + "production_corpus": { + "verified_sources": 0, + "independent_people": 0, + "independent_organizations": 0, + "source_types": [], + "roles": [], + "geographies": [], + "languages": [], + "books": 0, + "empirical_sources": 0, + "off_owned_source_share": null, + "vendor_source_share": null, + "consulting_source_share": null, + "propositions": 0, + "challenging_positions": 0, + "statements": 0, + "statement_verification_status": "no_production_statements", + "inventory_basis": [ + "data/sources contains README.md only", + "data/statements contains README.md only", + "data/propositions contains README.md only", + "data/people contains README.md only", + "data/dossiers contains README.md only", + "No production book records are present" + ] + }, + "full_topic_dossier_threshold": { + "verified_sources": 30, + "people": 15, + "organizations": 5, + "source_types": 3, + "independent_verification": true, + "human_approval": true + }, + "result": "FAIL", + "failed_requirements": [ + { + "metric": "verified_sources", + "actual": 0, + "required": 30, + "shortfall": 30 + }, + { + "metric": "people", + "actual": 0, + "required": 15, + "shortfall": 15 + }, + { + "metric": "organizations", + "actual": 0, + "required": 5, + "shortfall": 5 + }, + { + "metric": "source_types", + "actual": 0, + "required": 3, + "shortfall": 3 + }, + { + "metric": "independent_verification", + "actual": false, + "required": true + }, + { + "metric": "human_approval", + "actual": false, + "required": true + } + ], + "disposition": "research_gap_report_only", + "public_dossier_created": false, + "dossier_record_created": false, + "rationale": "The generator and dossier protocol permit verified production records only. All topic research beyond the discovery protocol remains in staging with human review pending.", + "unpublished_research_gap_context": { + "use_restriction": "Context for gap planning only; not dossier evidence.", + "source_records": 9, + "intellectual_sources_with_statements": 8, + "duplicate_renditions": 1, + "source_types": [ + "academic_paper", + "public_policy_document", + "research_report", + "working_paper" + ], + "distinct_named_authors_across_intellectual_sources": 68, + "enriched_person_records": 15, + "publishing_or_attributing_organization_groups": 9, + "languages": [ + "lang-en" + ], + "role_relevance_tags": [ + "role-ciso", + "role-cio", + "role-cto", + "role-general-counsel", + "role-board-director" + ], + "direct_role_specific_position_comparisons": 0, + "empirical_sources": 4, + "primary_synthetic_empirical_sources": 3, + "statements": 46, + "statements_human_review_status": "pending", + "candidate_only_proposition_ideas": 9, + "canonical_propositions": 0, + "formal_counterargument_statements": 0, + "off_owned_source_share": 0, + "vendor_involved_source_share": 0.25, + "consulting_source_share": 0, + "geography_note": "United States and global normative contexts plus synthetic benchmarks with no human geography; no region-comparative evidence." + } +} diff --git a/staging/batches/BATCH-2026-009/unresolved-questions.json b/staging/batches/BATCH-2026-009/unresolved-questions.json new file mode 100644 index 0000000..f27cdb1 --- /dev/null +++ b/staging/batches/BATCH-2026-009/unresolved-questions.json @@ -0,0 +1,85 @@ +{ + "batch_id": "BATCH-2026-009", + "questions": [ + { + "question_id": "UQ-001", + "question": "Which agent actions require a distinct runtime identity, and at what granularity should instances, sessions, and agent classes be distinguished?", + "domain": "Identity architecture", + "evidence_needed": "Interoperability tests and operator implementations", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-002", + "question": "How should authority attenuate across multi-agent delegation without breaking legitimate workflows?", + "domain": "Authorization", + "evidence_needed": "Deployment studies with measured task success and unauthorized-action rates", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-003", + "question": "Which credential lifetimes, rotation rules, and revocation mechanisms work at production agent speed and scale?", + "domain": "Lifecycle", + "evidence_needed": "Comparative implementation evidence", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-004", + "question": "What evidence should connect an agent action to a principal while preserving privacy and lawful anonymity?", + "domain": "Auditability and privacy", + "evidence_needed": "Legal analysis, civil-society positions, privacy engineering, and incident investigations", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-005", + "question": "How should legal responsibility be allocated when developer, deployer, user, platform, and tool provider share control?", + "domain": "Accountability", + "evidence_needed": "Multi-jurisdiction legal research and adjudicated or investigated cases", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-006", + "question": "How well do synthetic prompt-injection and vulnerability benchmarks predict production incidents?", + "domain": "External validity", + "evidence_needed": "Linked benchmark-to-production validation studies", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-007", + "question": "When does human approval improve control, and when does it create consent fatigue or ceremonial oversight?", + "domain": "Human oversight", + "evidence_needed": "Controlled user research and operational audit data", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-008", + "question": "Can automated monitors reliably oversee agents that act faster or at greater volume than human reviewers?", + "domain": "Monitoring", + "evidence_needed": "Independent evaluations, failure analyses, privacy and cost measurements", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-009", + "question": "Who owns agent identity governance across CISO, CIO, CTO, GC, CFO, CEO, and board roles?", + "domain": "Operating model", + "evidence_needed": "Role-specific interviews and documented decision-rights models", + "status": "unresolved", + "human_review_status": "pending" + }, + { + "question_id": "UQ-010", + "question": "How do requirements and implementations differ across industries and regulatory regions?", + "domain": "Industry and geography", + "evidence_needed": "Region- and sector-specific evidence with explicit denominators", + "status": "unresolved", + "human_review_status": "pending" + } + ] +} diff --git a/staging/batches/BATCH-2026-009/validation-results.md b/staging/batches/BATCH-2026-009/validation-results.md new file mode 100644 index 0000000..da9ac7f --- /dev/null +++ b/staging/batches/BATCH-2026-009/validation-results.md @@ -0,0 +1,35 @@ +# Validation results + +Overall: **PASS** + +| Check | Status | Detail | +|---|---|---| +| Structured YAML checks | PASS | Manifest and dossier protocol parse successfully. | +| Structured JSON checks | PASS | All machine-readable records parse and share the batch identifier. | +| Publication threshold checks | PASS | 0 verified production sources versus 30 required; all configured dossier requirements fail. | +| Public dossier withholding check | PASS | No thin public dossier or canonical dossier record was created. | +| Statement lineage checks | PASS | No staging statement is represented as production evidence. | +| Proposition lineage checks | PASS | No candidate proposition is represented as canonical. | +| Unsupported claim checks | PASS | The only editorial conclusion is the threshold failure and the report disclaims substantive answers. | +| Evidence-character checks | PASS | Potential empirical evidence is typed and its inference limit is explicit. | +| Source-concentration checks | PASS | Empty production denominator is null; staging vendor involvement is separately disclosed. | +| OFF ownership disclosure checks | PASS | Production denominator and staging share are separated; future OFF ingestion requires disclosure. | +| Role-representation checks | PASS | Role matrix prohibits inference from shared relevance tags. | +| Regional-representation checks | PASS | Region, study, organization, and regulatory coverage gaps are separated. | +| Consensus and AI-slop checks | PASS | Prohibited or generic language is absent and the editorial review records the check. | +| Near-duplicate content checks | PASS | 19 substantive paragraphs are unique. | +| Existing OFF content overlap checks | PASS | Three relevant OFF content families were reviewed for overlap only. | +| Claim-lineage checks | PASS | Ten gap-report claims have named lineage or calculation rows. | +| Staging isolation checks | PASS | The batch remains in staging and canonical dossier data is empty. | +| Human review checks | PASS | No machine-created record is marked human approved. | + +## Repository-wide checks + +Executed on 2026-08-15 after the dedicated batch checks: + +- Canonical data, provenance, review-status, publication, and content validations: PASS +- External-link audit: PASS +- Repository tests: 24 passed across 4 test files +- Public build: PASS; 28 static pages generated +- Search index and post-build artifact generation: PASS +- Git whitespace/error check: PASS