Human Conception Ledger as Provenance Infrastructure for Contested AI-Assisted Discovery: An Exemplary Use Case from the September 2026 Navier–Stokes and Fluid-Dynamics Priority Dispute

Abstract
Machine-readable watermarks and model-side provenance metadata answer one question: whether a generative system participated in a text, proof, or code artifact. They do not answer the questions that determine inventorship, academic credit, and institutional ownership: what a human conceived, contributed, decided, rejected, or directed, and when those acts occurred relative to later model use or later laboratory work. This note treats the publicly reported September 2026 dispute surrounding AI-assisted finite-time blowup results for forced incompressible fluid equations—and the later 8 September OpenAI institutional account of a multi-agent forced Navier–Stokes campaign—as an exemplary use case for the Human Conception Ledger (HCL). The mapping is evidentiary architecture, not adjudication. Public statements are treated as claimed events. The case is useful precisely because two asserted provenance histories can now be placed side by side while mathematical correctness, model participation, information flow, verification, and announcement sequence remain separate questions.

Figure 1. HCL mapping of two contested AI-assisted research timelines. Lane A maps the Buckmaster–Alpöge account; Lane B maps OpenAI’s 8 September afternoon institutional account. Dates and events are claimed events from primary public statements, not independent findings of fact.
1. Purpose and scope
This note does not determine whether any laboratory obtained a mathematically valid Navier–Stokes result, whether any result satisfies prize requirements, whether session or product-derived data influenced a later model, or whether any conversation was accurately recalled. Those questions belong to the parties, their institutions, formal verification, and eventual mathematical review. The narrower design question is: if Human Conception Ledgers had been running on both workflows from the start, which record types would now be inspectable, and which public ambiguities would be reduced to dated artifacts?
The companion introductory note argued that Anthropic-style watermarking and EU AI Act Art. 50(2) provenance marks travel with generated text but do not, on their own, confirm the full provenance of the content, because models are routinely used for proofreading, translation, rewriting, and completion of human-origin material (Chu, 2026a). The present case is that limitation under load: a Clay-adjacent research program, multi-model assistance, a personal collaboration that crosses competitor employers, rumor, and competing announcement incentives.

Figure 2. Retrospective narrative versus contemporaneous human-side provenance in AI-assisted discovery. Without structured provenance, questions of conception, research direction, model participation, verification, information flow, and authorship must be reconstructed from heterogeneous evidence after the fact. HCL and the Team Contribution Attribution Ledger (TCAL) instead link human insights and contribution events to AI Digital Witness records, verification and external corroboration, producing inspectable conception, contribution, and decision provenance. HCL records provenance; it does not determine the mathematical correctness of a result.
2. Publicly reported skeleton (claimed events only)
The chronology is now best read as two attributed lanes. Lane A is compiled from Buckmaster’s 8 September public statement, associated preprints and Lean repository, and related public materials concerning the Buckmaster–Alpöge collaboration. Lane B is compiled from OpenAI’s later 8 September public account of its institutional multi-agent campaign. A date appearing in one lane is not silently imputed to the other. Each line is a claimed event, not a verified finding.
Table T1. Dual claimed timeline (v1.1) — not a finding of fact. Each row aligns public claims by date and shows the HCL object that would have been generated if logged live.

● OpenAI institutional account (8 September afternoon). OpenAI dates relevant internal model training to 28 August; states that a multi-agent evaluation was launched on 1 September after rumors concerning Millennium-problem solves; reports roughly 10,000 agents over about 88 hours, including an unforced-Euler warmup involving roughly 100 agents and about 50 hours; dates arrival at a forced Navier–Stokes C/D candidate to 5 September; reports Lean formalization/verification using GPT-6 Astra with roughly 17 additional hours; and publicly announced the result on 8 September afternoon.
● Data-scope statement (OpenAI account). OpenAI’s public account addresses whether researchers or agents accessed the Buckmaster–Alpöge work before release and separately discusses residual product-derived data as a model-improvement question. For HCL purposes, this is best represented as an EEC/vendor attestation linked to PIIL session-scope and AI-DW provenance fields, rather than as a substitute for those fields.
● Prior human program. Forced-blowup constructions associated with Córdoba and Martínez-Zoroa are identified by Buckmaster as the originating research program; LLM work is described as pushing that program from rough forcing toward smooth forcing and incompressible Euler.
● Collaboration form. Described as a personal collaboration between Tristan Buckmaster (NYU Courant) and Levent Alpöge (employed at Anthropic), with no institutional agreement; tools stated to be paid from personal research funds, including OpenAI usage.
● Models used (claimed). Anthropic Claude; OpenAI Codex / GPT-class systems (including a “GPT-5 6 Sol” designation in the statement); later “Astra” limited to writeups and auditing.
● Human-dated milestones (claimed). Slow literature and preliminary upgrades over about one year; finite-time blowup with smooth forcing for IPM earlier; Boussinesq and Euler with smooth forcing obtained about 15 August 2026; first LLM-generated proof described as nearly unreadable; Lean verification about 22 August 2026; subsequent around-the-clock human effort to make arguments readable.
● What was released. Three results made public: finite-time blowup with smooth forcing for incompressible porous media, Boussinesq, and 3D incompressible Euler. Hypo-dissipative Navier–Stokes mentioned as believed-in-hand but withheld pending finished Lean verification and a presentable writeup.
● Quality admission. Buckmaster states the Euler writeup in particular “can only be described as AI slop,” and that the preferred scientific norm would have been weeks of human rewriting before first public text.
● Rumor and outreach. On 3 September 2026, with rumors that Anthropic had resolved a major open problem and tips that progress information had reached OpenAI, Buckmaster emailed a mathematician affiliated with OpenAI to pre-empt misattribution and to state the personal (non-institutional) character of the work.
● Subsequent calls (claimed). 6 September 2026 conversations including Sébastien Bubeck. Claimed contents include: an internal OpenAI model described as producing an approximately 100-page proof of forced Navier–Stokes blowup (Fefferman options c/d; smooth forcing); initial characterization of “very little human input” later complicated by description of a team, staged easier problems including Euler, Codex-written prompts, and large compute; first prompt timing eventually placed in “the past few days,” after information about the Buckmaster–Alpöge work had reached OpenAI; unanswered or incomplete answers on training or access to the pair’s Codex sessions.
● Authorship and announcement proposals (claimed). Coordinated consecutive posting (Euler then OpenAI Navier–Stokes), or Buckmaster alone writing up a Navier–Stokes result crediting an internal OpenAI model, with repeated suggestion that Alpöge be omitted because of Anthropic employment. Both declined. Further remarks about career consequences are alleged and subsequently denied as “false and inflammatory.”
● Parallel independent line. Anandkumar and collaborators publicly described a stable singularity construction for 3D Euler on R³ without forcing, via physics-informed neural networks plus stability analysis—methodologically distinct, same news cycle.
Later on 8 September, OpenAI publicly announced a forced Navier–Stokes result that it described as resolving alternatives C and D of the official formulation, together with a Lean formalization, while separately stating that it was not making a Clay Prize claim. The present note does not independently assess the mathematical correctness, interpretation, priority, or prize eligibility of that result. That separation—mathematical claim, verification artifact, provenance, and prize status—is itself part of the HCL design problem.
3. Why watermarking is the wrong instrument
If every LLM-touched paragraph or proof artifact carried a Claude, Codex, or OpenAI watermark, the public would still not know: (i) whether the forced-smooth strategy was inherited from prior human work or proposed by a model; (ii) which humans selected intermediate problems and scaled an agent campaign; (iii) how first-prompt and model-development dates relate to rumor or information-flow events; (iv) whether product-derived or session-derived data were in scope for model improvement or retrieval; (v) who directed formal verification; (vi) which model produced which candidate artifact; or (vii) who made release, withholding, authorship, and announcement decisions. These are human conception, contribution, boundary, and decision events. They are the objects HCL records.
4. HCL record types used in the mapping
Notation follows the HCL specification: Human Contribution Events (HCE); Human Insight Nodes (HIN); AI Digital Witness (AI-DW); Human Event Verification Bundles (HEVB); Corroborated Human Event Bundles (CHEB); External Event Corroboration (EEC); Private Immutable Inventor Ledger (PIIL). Team-level extension, where two or more humans plus one or more models act, is the Team Contribution Attribution Ledger (TCAL).
5. Worked mapping table
Each row takes a dispute surface that is now being argued in public and states (a) the HCL object that should have existed, (b) the minimum fields that object carries, and (c) what would now be checkable instead of narrated.
Dispute surface (as now argued in public) | HCL / TCAL object | Minimum fields the ledger would have carried | What would now be inspectable |
Origin of the research program: Córdoba–Martínez-Zoroa forced blowup vs “the model solved it.” | HIN + HCE (prior art / adoption) + EEC | Dated human statement of the inherited program; citation handles; explicit “we adopt X, we do not claim X”; hash of the papers used as starting point. | A third party could see that the basic line of attack is attributed to named prior humans before any LLM session, blocking prize-language collapse into “Claude solved NS.” |
Personal collaboration vs Anthropic or NYU project. | HCE (boundary) + PIIL governance fields + TCAL party table | Party identifiers; employer; “institutional agreement: none”; funding source for compute; license of outputs; who may announce. | Employment at Anthropic would be a recorded affiliation, not a reason invented after the fact to drop an author. Ownership would be a field, not a negotiation after rumor. |
Who had the first non-obvious human insight that smooth forcing + this construction reaches Euler / Boussinesq. | HIN (conception) with pre-model timestamp | Natural-language insight; why it was not in the prior paper; human author of the insight; time; optional sketch; “model not yet queried on this node.” | Conception date would be independent of the first successful LLM proof dump (~August) and independent of September announcement politics. |
Year of slow literature work vs one-month “Deep Blue” compression story. | HCE series (study / upgrade / dead-end) | Chronological HCEs: papers read, lemmas upgraded, failed routes, IPM milestone date. Each hashed. Negative results kept. | The public narrative could not flatten a year of human direction into a single model session. Duration of human attention would be visible. |
Which model did which job (search, proof candidate, rewrite, audit). | AI-DW per session, linked to parent HCE/HIN | Vendor/model/version; purpose tag {search | candidate-proof | formalize | rewrite | audit}; prompt hash; output hash; human accept/reject/edit; compute spend if known. | Claude vs Codex vs Astra would not be a single “AI helped” blob. Astra limited to writeup/audit would be a typed witness, not a rumor. |
First LLM-generated proof “the most horrendous I have ever read.” | AI-DW (candidate-proof) + HCE (human rejection / rewrite directive) | Raw model output hash; human verdict = reject-as-unreadable; directive to formalize in Lean rather than publish prose; date ~pre-22 Aug. | Establishes that model output was not treated as a paper. Human refusal is a first-class event, which watermarking never records. |
Lean verification on 22 August as the actual confidence event. | HEVB (formal check) + EEC (Lean repo commit) | Checker, commit hash, statement checked, who drove the formalization, who signed off that informal and formal objects match. | Confidence would attach to a dated formal artifact, not to a social-media prediction that “Claude solved Navier–Stokes.” |
Decision to withhold hypo-dissipative Navier–Stokes until Lean finishes and a writeup exists. | HCE (decision / non-release) | Decision owner; reason codes {verification-incomplete, presentation-inadequate}; artifact embargo flag in PIIL. | Withholding would read as scientific control, not as hiding a Millennium result. |
Admission that public Euler text is “AI slop” released under pressure. | HCE (release decision) + AI-DW (rewrite) + quality flag | Intended hold period; actual release time; stated external pressure; human sign-off that the text does not yet meet community standard. | Priority and presentation quality would be separable. A ledger can timestamp a result without forcing an unreadable first public draft. |
3 September rumor that Anthropic resolved a major open problem. | EEC (rumor capture) + HCE (correction email) | Hash of rumor as encountered; text of Buckmaster email to OpenAI-affiliated mathematician; time sent; stated facts (personal collab, no institutional effort). | The correction email becomes an immutable human decision event: pre-emptive clarification, not “we were hiding it.” |
Claim that information about progress reached OpenAI before OpenAI’s first prompt. | EEC (information-flow) + AI-DW (first foreign prompt, if they also logged) | What was transmitted, by whom, to whom, when — originating side. Receiving lab HCL: first prompt time, prompt hash, whether the prompt was model-written. | “First prompt in the past few days” would be a field, not a reluctant concession on a call. Independent ledgers could be compared without trusting either narrative alone. |
“Very little human input” vs later description of a team, staged Euler warmup, Codex-authored prompt, large compute. | TCAL party table + AI-DW chain + HCE (prompt authorship) | Human operators; intermediate problems assigned; whether the shown prompt was generated by Codex; compute envelope; “human input” ordinal {none, light direction, staged program}. | The phrase “very little human input” would be unsupported as a summary. The ledger would show a staged human research program or its absence. |
Did an internal model train on or retrieve the pair’s Codex drafts? | AI-DW access log + PIIL session-scope field + EEC (vendor attestation) | Session IDs of the pair’s Codex use; vendor attestation {no user-data lookup / training exclusion / unknown}; time of attestation. | The question Buckmaster says was not cleanly answered becomes a binary or ternary field with a named attestor. Silence is itself recorded. |
Proposal that Alpöge be removed because he works at Anthropic. | HCE (authorship proposal) + TCAL contribution matrix + HCE (refusal) | Who proposed omission; stated reason; contribution matrix already showing Alpöge’s HCEs/HINs; refusal by Buckmaster; time. | Author-dropping would appear as a dated proposal with a recorded motive. Contribution would already be visible enough to make the proposal evaluable. |
Proposal: consecutive announcements, or Buckmaster solo-writes OpenAI’s NS result crediting “an internal model.” | HCE (offer) + HCE (decline) + optional joint CHEB | Exact terms offered; who was on the call; decline text; whether a joint CHEB was even possible given an unseen proof. | Announcement choreography would be logged as negotiation, not as the natural order of discovery. An unseen 100-page proof would have no HIN/HEVB on Buckmaster’s ledger. |
Alleged career remarks vs later public denial. | EEC (contemporaneous note) — optional, high-sensitivity | If either party ran HCL, a contemporaneous memo-to-ledger (hash of notes written immediately after the call) plus any recording policy. | He-said would still exist, but each side would have a same-day hashed note. Weaker than audio; stronger than a week-later statement. |
Anandkumar et al. unforced Euler via PINN + stability, announced in the same window. | Separate PIIL / independent HIN tree + EEC (method contrast) | Method tag {PINN-then-stability} vs {LLM-directed proof search on forced constructions}; no-forcing vs smooth-forcing; start date of that program. | Two independent human conception trees would be visible. Neighboring problems and different AI modalities, not a single stolen result. |
Social prediction “Claude has solved Navier–Stokes… before the IPO.” | Out of band (press) but HCL-relevant as EEC pollution | HCL cannot stop prediction posts. It can publish a short public CHEB: problem identifier, forced vs unforced, verified vs withheld, models as witnesses not inventors. | Prize-language inflation would still occur, but the authors would have a one-page machine-readable correction already hashed. |
6. What an HCL packet for this collaboration would have looked like
Minimally, the pair’s PIIL would contain four layers, exportable as a Corroborated Human Event Bundle if they later chose to disclose:
● Layer A — Conception tree. HIN for adoption of the Córdoba–Martínez-Zoroa program; HINs for each equation class (IPM, Boussinesq, Euler, withheld hypo-dissipative NS); explicit non-claim on unforced Clay Navier–Stokes.
● Layer B — Direction log. HCEs for human choices that models cannot own under current patent and academic norms: choice of smooth forcing, choice to formalize before polishing prose, choice to withhold NS, choice to email OpenAI on 3 September, choice to decline authorship edits on 6 September.
● Layer C — Witness log. AI-DW objects for Claude / Codex / Astra with purpose tags. Rejected drafts kept. Lean commits attached as HEVBs. No model listed as inventor.
● Layer D — Boundary log. TCAL party table (two humans, two employers, no joint venture); compute invoices; “no institutional agreement”; announcement rights. This layer is what makes “drop the Anthropic employee” a visible rewrite rather than a quiet edit.
A receiving laboratory that also ran HCL would add its own Layer C: model-genesis records, first prompt time, prompt authorship, agent orchestration, compute, team roster, candidate-proof hashes, and formal-verification artifacts; and its own Layer D: institutional boundaries and whether user-session or product-derived data were in scope. OpenAI’s later public account supplies an asserted institutional chronology for several of these fields. Comparing two CHEBs is how multi-lab priority should be discussed. Comparing press releases is only a proxy for that comparison.

Figure 3. Conceptual architecture of an HCL provenance packet for AI-augmented collaborative research. The Private Immutable Inventor Ledger (PIIL) organizes provenance into four linked layers: a human conception tree, human direction log, AI and verification witness log, and team/governance boundary log. Selected records may be assembled into a Corroborated Human Event Bundle (CHEB) for disclosure to referees, patent examiners, institutions, or collaborating or competing research groups while the underlying research ledger remains private. AI systems are represented as digital witnesses rather than substitutes for separately attributable human conception, contribution, and decision-making.
7. What HCL does not do
A ledger does not make a proof correct. Lean can formally verify a specified formal object; mathematical review still determines whether the formalized statement and its interpretation establish the claimed result. A ledger does not make people honest; it makes inconsistency more inspectable. It does not assign a Fields Medal or Clay Prize, and it does not determine whether a forced Navier–Stokes result satisfies any prize requirement. HCL records provenance; it does not adjudicate mathematical correctness, prize eligibility, or priority by itself.
8. Design implication
The episode is being received as a story about laboratory rivalry. Structurally it is a story about missing human-side provenance. Watermarks scaled; conception records did not. As soon as multiple capable teams can aim models or large agent populations at the same open problem in days, announcement order, prompt time, model genesis, session isolation, information flow, human direction, and coauthor identity become first-order scientific infrastructure. HCL was specified for that infrastructure: capture human-origin inventive acts, treat models as digital witnesses, and emit selectively disclosable bundles.
The subsequent OpenAI disclosure illustrates the point directly. Once a second institution supplies its own dates, model-development history, agent chronology, verification history, and account of information flow, the provenance problem becomes a comparison between two asserted event histories. HCL is designed to make such histories contemporaneously recordable and selectively corroboratable rather than retrospectively reconstructed from competing public statements. The speed demonstrated by human-plus-model and multi-agent workflows is the reason human conception, contribution, direction, verification, and release decisions must become more visible, not less.
9. Status of this note
This is an exemplary mapping for discussion and Zenodo companion use. It is not legal advice, not an independent finding that any mathematical claim is correct, not a claim that any party did or did not train on or retrieve another party’s sessions, and not an endorsement of any prize claim. Facts should be updated if primary statements are corrected. Version 1.1 adds OpenAI’s 8 September institutional account as a second claimed timeline and uses it to demonstrate cross-ledger comparison.
References
Buckmaster, T. (2026). Public statement and linked manuscripts (Euler, IPM, Boussinesq) and Lean repository. Courant Institute pages dated 8 September 2026.
Chu, M. B. (2026a). Preserving Human Conception, Contribution, and Decision-Making in the Age of AI: Introduction to the Human Conception Ledger for Human Provenance. Zenodo. https://zenodo.org/records/21880405
Chu, M. B. (2026b). Human Conception Ledger (HCL): A Framework for Provenance, Attribution, and Human Inventorship in AI-Augmented Systems. Zenodo. https://zenodo.org/records/19198704
Anandkumar, A., et al. (2026). Stable singularity of the Euler equations on R³ without forcing. https://anima-ai.org/2026/09/07/stable-singularity-of-the-euler-equations-on-r3-without-forcing/
Fefferman, C. L. Official Clay Mathematics Institute statement of the Navier–Stokes existence and smoothness problem (options including external force).
Public posts and reporting of 8 September 2026 summarizing the above primary statements (used only to locate claimed events; not treated as independent authority).
OpenAI. (2026). Public research note and X thread describing a multi-agent forced Navier–Stokes campaign, model-development chronology, Lean verification, concurrent work, and data-scope statement. 8 September 2026. https://x.com/openai/status/2097374640582668336
This paper is also available on Zenodo. v1
This version v1.1 to be uploaded when Zenodo up and running again.



Comments