Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

WG3 — SCE Working Group on AI Assistance, Reproducibility & Standards

Part of the Research Program for AI as Babel Fish for Structural Economics.

Chair: Alan Lujan (designated June 2026).

Scope

This working group addresses assisted use, institutional uptake, and sustainability: AI tools for model authoring and interpretation, grounded in the formal semantic core; standards for reproducibility and robustness; and governance of the language as a community standard.

Aims

Develop symbolically grounded AI tooling — in which a language model supplies informal reasoning but must satisfy the typed Bellman-core and formal knowledge base of WG1 — for verified translation from informal model descriptions to formal modular DP and valid language specifications. In parallel, develop the reproducibility infrastructure (automated robustness testing, referee/data-editor workflows, REMARK-style archives) and the governance arrangements required for the language to function as an accepted standard in the profession.

Scientific rationale

AI-assisted modeling is scientifically useful only if outputs are constrained by formal semantics rather than by prose pattern-matching. The symbolically grounded approach now emerging in AI-for-science (see prior art §5) treats a symbolic semantic layer as a hard constraint on neural components. Separately, a specification language achieves little without adoption: reproducibility standards, systematic robustness analysis, and transparent governance are the institutional mechanisms through which technical capability translates into professional practice. Together, these address the question of whether a published result depends on undocumented computational choices.

Current state

AI assistance (from matsya)

Reproducibility and standards

Cross-subfield translation and reproducibility

Two of WG3’s capabilities extend naturally across economics subfields (CGE, ABM, applied IO):

The representational and numerical halves of cross-subfield integration sit in WG1 and WG2. See prior art §6 for cross-field precedents (OpenFermion as a translation bridge; the standing consortia behind OPTIMADE and ESMF).

Five-year milestones

Open research questions

Illustrative project proposals

Develop verified translation: a paper→formal-MDP→language pipeline in which every AI-produced specification is checked against the Bellman-core type system and the formal knowledge base, evaluated on the EconDP benchmark of published structural models.

Pilot the referee workflow: in partnership with a journal data editor, express a set of recently published structural results in the language, run automated robustness tests over their declared numerical methods, and report which conclusions are sensitive to computational choices.

Indicators of progress

Active Matsya users and sessions; translation accuracy (parses, round-trips, verification against the formal knowledge base); EconDP benchmark scores; REMARK reproductions in the language; journals and data editors piloting the workflow; robustness reports produced; governance documents ratified; sustained funding secured.

Potential funding sources

JHU DS/AI Institute, NSF (CISE / robust AI for science; open science / replication / cyberinfrastructure), central banks (robust policy models), FRO/Convergent, philanthropic follow-on.

Foundations

Sargent–Stachurski (the DP-theory corpus Matsya grounds in); the Bellman calculus and formal knowledge base (WG1) as the symbolic substrate; symbolically grounded AI (the broader paradigm, a.k.a. neuro-symbolic); BufferStockTheory (canonical examples); the Econ-ARK REMARK standard; Backus FFP (the meaning-vs-representation split that makes robustness testing well-defined); the econ-ark.org reproducibility mission.