Tab 02 · Evidence Collection · 5 to 18 October 2026 · Lead: Rocky Hill

Evidence Collection

A two week window to gather the material the project will be built on. This tab holds only NEW evidence and leads identified after the Project Plan was drafted; the foundation sources stay on the Initial Research tab untouched. Everything here is aimed at the specific evidence gaps named in Project Plan section 3.4.
ExpandGo beyond what the team already has. Find more sources: field experiments, standards, non Western data, DXC materials, case studies, and interviews.
AssessCheck each source for quality before it enters the base. Is it credible, current, relevant, and free of unexamined commercial interest?
OrganizeSort the evidence by project objective so it is clear which objectives are well supported and which still need more.

Expand: New Source Leads, Grouped by the Gap They Close

Project Plan section 3.4 named three evidence gaps: no consistent definition or measure of AI value, little evidence directly linking trust and change readiness to financial results, and an evidence base dominated by Western consultancy research. The groups below attack those gaps, plus two more needs that surfaced in scoping: governance benchmarks for the maturity model, and fresher failure data. Each card states why it earns a place now.

Gap 1 · Causal, measured evidence of AI value (not self reported surveys)

The strongest weakness in the current base: almost every value number is a self reported executive survey. These are randomized and natural field experiments with measured outputs, the highest quality evidence available on what AI actually does to productivity and quality.

Organization Science · 2026

Navigating the Jagged Technological Frontier

Peer Reviewed Journal

Dell'Acqua, McFowland, Mollick, Lifshitz, Kellogg and colleagues ran a field experiment with 758 BCG consultants. Inside AI's capability frontier, consultants with AI finished more tasks, faster, at measurably higher quality; outside the frontier, AI use made results worse. Originally the famous 2023 HBS working paper, now published in a top peer reviewed journal.

Why now: Closes the measurement gap with experimental data, and the jagged frontier concept gives Q2 (What AI Is Not) an academic backbone to pair with Gartner's agent washing.
Open Source
The Quarterly Journal of Economics · 2025

Generative AI at Work

Peer Reviewed Journal

Brynjolfsson, Li, and Raymond studied 5,172 customer support agents in a staggered real world rollout. AI assistance raised productivity 15% on average, with the largest gains for less experienced workers, improved customer sentiment, and improved employee retention.

Why now: The retention finding is the closest thing in the literature to a measured link between worker experience and business results, exactly the trust to outcomes connection gap 2 of the plan flags as missing.
Open Source
Science · 2023

Experimental Evidence on the Productivity Effects of Generative AI

Peer Reviewed Journal

Noy and Zhang's randomized experiment on professional writing tasks: ChatGPT cut time taken by about 40% while raising output quality, and it compressed the gap between weaker and stronger performers rather than widening it.

Why now: A second independent experimental result confirming where value lands (task productivity) and for whom (lower skilled workers most), published in one of the two most prestigious journals in science.
Open Source
Harvard Business School · 2025

The Cybernetic Teammate

Academic Working Paper

Dell'Acqua, Sadun, Lakhani and colleagues' follow up field experiment at Procter and Gamble: individuals working with AI matched the performance of two person teams without AI, and AI use broke down the usual silo between commercial and technical staff.

Why now: Moves the value question from individual productivity to team and operating model design, the exact layer where McKinsey says value is won or lost.
Open Source

Gap 2 · Linking trust, readiness, and change to outcomes

The plan concedes that trust evidence measures perception, not results. These sources get closer to the behavior and leadership side of the link.

McKinsey · Jan 2025

Superagency in the Workplace

Consultancy Research

Survey finding that employees are far more ready for AI than leaders believe: employees are already using AI for a large share of their work while leaders underestimate usage by roughly three times, and the biggest barrier to scaling is leadership, not employee resistance.

Why now: Flips the change management assumption in the plan. If resistance sits in the leadership layer, the maturity model's readiness dimension needs to measure leaders, not just the workforce.
Open Source
Microsoft WorkLab · Annual

Work Trend Index

Vendor Research

Microsoft and LinkedIn's annual study of tens of thousands of workers across 31 countries, tracking real usage telemetry alongside survey data: employee adoption running ahead of official programs, the rise of human agent teams, and the emergence of frontier firms.

Why now: One of the few large datasets pairing stated attitudes with actual usage behavior, which strengthens the behavioral adoption versus technical implementation question from the sponsor meeting.
Open Source

Gap 3 · Widening the lens beyond Western consultancy data

The plan names the Western, consultancy heavy skew as a limitation. These intergovernmental datasets cover most of the world with published methods and no product to sell.

International Monetary Fund

AI Preparedness Index

Intergovernmental Dataset

IMF index covering 174 economies across digital infrastructure, human capital and labor policies, innovation, and regulation. Shows advanced economies better positioned but more exposed to disruption, while emerging markets face a widening readiness gap.

Why now: Country level, method published, free of commercial interest. The backbone for the international comparison the RFS requires, beyond the adoption percentages already carded.
Open Source
OECD

OECD.AI Policy Observatory

Intergovernmental Observatory

Live database of national AI strategies, policies, and incidents across 70 plus jurisdictions, plus the OECD AI Principles that most national frameworks trace back to. Includes trend data on investment, skills, and regulation by country.

Why now: Lets the team evidence the regulatory divergence claims (Europe regulation first, North America innovation first, APAC scale first) with primary policy data instead of secondhand summaries.
Open Source

Gap 4 · Governance benchmarks the maturity model can anchor to

Dr. Kokkonen asked for a maturity model with stage gates. These three are the recognized external benchmarks a governance dimension can be graded against, which keeps the framework from being arbitrary.

NIST · United States

AI Risk Management Framework (AI RMF)

Government Framework

The US national framework for trustworthy AI: Govern, Map, Measure, Manage. Voluntary, widely adopted by enterprises as the de facto governance baseline, with a Generative AI profile added for GenAI specific risks.

Why now: A ready made, citable structure for the governance dimension of the diagnostic. An organization's distance from AI RMF practices is a measurable maturity gap.
Open Source
ISO/IEC · International

ISO/IEC 42001:2023 AI Management Systems

International Standard

The first certifiable international standard for AI management systems, covering governance, risk, lifecycle management, and continual improvement. Organizations can now be audited and certified against it.

Why now: Certification is a hard, binary stage gate, exactly the graded timeline artifact the sponsor asked for. Few maturity models in the consultancy literature anchor to it yet, which is a differentiation opportunity.
Open Source
European Commission

The EU AI Act: Regulatory Framework on AI

Regulation

The world's first comprehensive AI law, phasing in through 2026 and 2027 with risk tiered obligations and significant penalties. The reference point for Europe's regulation first posture and for compliance as a value protection activity.

Why now: Turns the international perspective from description into consequence: the same AI investment carries different compliance cost and risk by region, which belongs in the framework's regional factor.
Open Source

Gap 5 · Fresher failure and abandonment data

The initial base leans on the contested MIT 95% figure and Gartner's forward prediction. These add measured abandonment rates and systematically studied root causes.

RAND Corporation · 2024

The Root Causes of Failure for AI Projects

Policy Research Institute

RAND interviewed 65 experienced data scientists and engineers and synthesized five root causes of AI project failure, led by leaders misunderstanding or miscommunicating the problem to be solved, with inadequate data and infrastructure close behind. Technology shortcomings rank last.

Why now: Independent, non commercial confirmation of the organizational over technical thesis, insulating the Q4 argument from the consultancy bias critique in plan section 3.4.
Open Source
S&P Global Market Intelligence · Mar 2025

AI Abandonment Jumps from 17% to 42%

Media Summary of Survey

S&P's survey of 1,000 plus enterprises across North America and Europe: the share of companies abandoning most of their AI initiatives jumped from 17% in 2024 to 42% in 2025, with an average 46% of proofs of concept scrapped before production. Cost, data privacy, and security cited as top causes.

Why now: A measured abandonment rate, not a prediction, and a dramatic one year trend line for the report's problem statement. Chase the underlying S&P report for the primary citation.
Open Source
Gartner · July 2024

30% of GenAI Projects Abandoned After Proof of Concept

Analyst Research

The predecessor to the agentic prediction already carded: Gartner forecast at least 30% of generative AI projects would be abandoned after proof of concept by end of 2025, citing poor data quality, inadequate risk controls, escalating costs, and unclear business value.

Why now: Pairing the 2024 prediction with S&P's 2025 measured 42% shows the forecast was conservative, a clean before and after exhibit for the report.
Open Source

Primary evidence the team can create

Expanding is not just finding documents. The team's own position gives it primary evidence no consultancy has, and Ducere explicitly invites surveys and personas as deliverables. Anything involving people goes through the plan's section 5 ethics and approval gate first.

Team Quantum · Primary Research

Practitioner Interview Program

Primary Evidence

Six to ten short semi structured interviews with leaders and practitioners across the team's networks, using the LinkedIn channel Dr. Kokkonen opened (she accepts all connection invites and encouraged cross group networking). Target a spread of industries, regions, and seniority, asking how they define AI value, where it leaked, and what they would measure differently.

Why now: Directly fills the trust to outcomes gap with firsthand accounts, and gives the report quotable, original material no other group will have. Requires the ethics and consent check in plan section 5 before any interview begins.
No external link. Owner and schedule to be set in the team's next working session.
Team Quantum · Primary Research

Stakeholder Perception Pulse Survey

Primary Evidence

A short anonymous survey built on the Melbourne and KPMG question structure, distributed through the team's combined networks across the six RFS stakeholder groups. Even 50 to 100 responses gives the perception heat map an original data layer to compare against the global benchmarks.

Why now: Ducere's deliverable guidance explicitly invites survey creation, analysis, and synthesis. Reusing validated question wording keeps it defensible. Ethics gate applies.
No external link. Draft instrument before 18 Oct so distribution can run during analysis.
Team Quantum · Desk Sweep

DXC Public Materials Sweep

Primary Evidence

A systematic pass through DXC's own public record: the full AdvisoryX study download, annual report and investor materials (how DXC describes AI demand to shareholders), newsroom releases, and Dr. Kokkonen's publication and speaking trail including her Trust in Government's use of AI talk.

Why now: Client data is off limits, but DXC's public voice is fair game and shows the sponsor's worldview in its own words, keeping the framework aligned without confirmation bias (test it against the independent sources).
Start at dxc.com/newsroom and dxc.com/investor-relations; log each find in the shared tracker.
Team Quantum · Internal

Cross Industry Mini Cases from the Team Itself

Primary Evidence

Five members spanning software sales (UK and global), digital marketing operations (London), government contracting operations (Hawaii), and construction and refrigeration services (Dallas). Each member writes a one page structured vignette of AI value creation or value loss observed in their own sector, using a common template.

Why now: Dr. Kokkonen called the team's neutrality and cross industry experience a strength. Structured vignettes convert that from a nice line in the intro into actual evidence, clearly labeled as practitioner observation.
No external link. Common template to be agreed before writing so the vignettes are comparable.

Assess: The Quality Screen Every Source Passes Before It Enters the Base

Run each new source through the six checks below before it is added to a tab. A source can still be used if it fails a check, but the weakness gets written down next to it, the way section 3.1 of the Project Plan already does with its Critical Consideration column. The standing rule from section 3.4: consultancy and vendor claims are triangulated against academic sources before they carry weight in findings.

CheckQuestion to AskPasses WhenRed Flags
CredibilityWho produced it and what is their track record?Named authors or established institution; peer review or published editorial standards.Anonymous content, AI generated blog farms, no institutional home.
CurrencyIs it current enough for a field moving this fast?2024 or later for market data; older is fine for theory (J curve, trust theory) where the concept, not the number, is being used.Pre 2023 adoption or ROI statistics presented as current.
RelevanceWhich research question or objective does it serve?Maps cleanly to at least one of Q1 to Q4, the international lens, or a maturity model dimension.Interesting but unmappable; collect somewhere else, not in the base.
Method transparencyCan you tell how the numbers were produced?Sample size, population, and method published; limitations acknowledged.Headline percentages with no methodology anywhere (the DXC AdvisoryX study itself carries this caution).
Commercial interestDoes the publisher sell a remedy for the problem it describes?No stake, or the stake is declared and the claim is triangulated against an independent source.Vendor research used alone to justify a finding that benefits the vendor.
Primary vs secondaryIs this the original source or a retelling?Original report, paper, or dataset cited; media retellings used only for critique and color.Citing a news summary when the underlying report is freely available.

Organize: Evidence Coverage Against the Project's Needs

Objectives O1 to O5 are still being finalized in section 1.5 of the Project Plan, so coverage is organized against the four RFS research questions plus the three framework needs the sponsor set. When the objectives are locked, this table gets recut against them. Status reflects the base as of 3 October: the job of this phase is to move every row to Strong by 18 October.

Project NeedStatusWhat Is Already StrongGap to Close by 18 OctNew Leads on This Page
Q1 What is AI valueAdequateSurvey evidence (BCG, McKinsey, Deloitte) and the academic definition (Enholm).Causal, measured evidence; a working definition tested against it.Jagged Frontier, Generative AI at Work, Noy and Zhang, Cybernetic Teammate.
Q2 What AI is notStrongGartner agent washing, AI theater, hype critique both directions.Academic grounding for capability limits.Jagged frontier concept (capability is uneven, not general).
Q3 How AI is perceivedAdequateMelbourne and KPMG global data, Stanford perception gap, trust theory (Glikson and Woolley).Perception linked to behavior and outcomes; original stakeholder data.Superagency, Work Trend Index, pulse survey, practitioner interviews.
Q4 Why businesses struggleStrongDXC five barriers, MIT learning gap, workflow redesign evidence, CEO testimony.Non commercial confirmation; measured (not predicted) failure rates.RAND root causes, S&P 42% abandonment, Gartner 2024 baseline.
Value over time horizonsAdequateJ curve theory, Deloitte one to five year agentic expectations.Firm level evidence of the lag and recovery.Field experiments above give the early, task level end of the curve.
International perspectiveNeeds MoreAdoption and trust percentages by country.Non Western, non consultancy data; regulation as a regional cost factor.IMF Preparedness Index, OECD.AI Observatory, EU AI Act.
Maturity model and diagnostic inputsNeeds MoreDXC five barriers and Enholm enablers as candidate dimensions.External benchmarks so maturity grades are defensible, not invented.NIST AI RMF, ISO/IEC 42001, EU AI Act risk tiers, Edge Framework phases.

New References Added This Phase (Harvard Style)

Only the sources introduced on this tab. The foundation references stay on the Initial Research tab. Verify dates and details against each original publication before final submission.

The S&P Global entry cites the news summary; replace it with the underlying S&P Global Market Intelligence report citation once the team retrieves the primary document.

Brynjolfsson, E., Li, D. and Raymond, L. (2025) 'Generative AI at Work', The Quarterly Journal of Economics, 140(2), pp. 889 to 942. doi: 10.1093/qje/qjae044.

Dell'Acqua, F., McFowland, E., Mollick, E., Lifshitz, H., Kellogg, K.C., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K.R. (2026) 'Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality', Organization Science. doi: 10.1287/orsc.2025.21838.

European Commission (2024) Regulatory Framework on Artificial Intelligence. Available at: https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai (Accessed: 3 October 2026).

Gartner (2024) Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025 [Press release]. Available at: https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025 (Accessed: 3 October 2026).

Harvard Business School AI Institute (2025) The Cybernetic Teammate: A Field Experiment on Generative AI Reshaping Teamwork and Expertise. Available at: https://aiinstitute.hbs.edu/?p=26305 (Accessed: 3 October 2026).

International Monetary Fund (2024) AI Preparedness Index. Available at: https://www.imf.org/external/datamapper/datasets/AIPI (Accessed: 3 October 2026).

International Organization for Standardization (2023) ISO/IEC 42001:2023 Information Technology, Artificial Intelligence, Management System. Available at: https://www.iso.org/standard/81230.html (Accessed: 3 October 2026).

McKinsey & Company (2025) Superagency in the Workplace: Empowering People to Unlock AI's Full Potential. Available at: https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/superagency-in-the-workplace-empowering-people-to-unlock-ais-full-potential-at-work (Accessed: 3 October 2026).

Microsoft (2025) Work Trend Index Annual Report. Available at: https://www.microsoft.com/en-us/worklab/work-trend-index (Accessed: 3 October 2026).

National Institute of Standards and Technology (2023) AI Risk Management Framework (AI RMF 1.0). Available at: https://www.nist.gov/itl/ai-risk-management-framework (Accessed: 3 October 2026).

Noy, S. and Zhang, W. (2023) 'Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence', Science, 381(6654), pp. 187 to 192. doi: 10.1126/science.adh2586.

OECD (2026) OECD.AI Policy Observatory. Available at: https://oecd.ai (Accessed: 3 October 2026).

RAND Corporation (2024) The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed. Available at: https://www.rand.org/pubs/research_reports/RRA2680-1.html (Accessed: 3 October 2026).

This Week Health (2025) AI Project Failures Surge to 42% as Companies Struggle to Scale [Summary of S&P Global Market Intelligence survey]. Available at: https://thisweekhealth.com/news_story/ai-project-failures-surge-to-42-as-companies-struggle-to-scale (Accessed: 3 October 2026).