Back
Agentic problem solving / Findings of EMNLP 2025

ModelingAgent

Bridging LLMs and real-world mathematical modeling.

A tool-using modeling workflow that develops ideas, gathers data, computes models, and refines a report through critique. ModelingBench supplies the tasks; ModelingJudge evaluates the results.

ModelingAgentCONCEPT / MODELING WORKFLOW
FROM QUESTION TO MODELING REPORTMODELINGAGENTTool use · retained context · iterative refinement01ProposeModeling ideas02Find dataGround the model03ModelFormulate & compute04WriteExplain the resultsCritique → refine → continue
One modeling workflow · tools · iterative refinement
CONCEPTUAL RESEARCH COVER / ModelingAgent
THE RESEARCH QUESTION

How can a modeling workflow turn open-ended problems into defensible mathematical models?

Mathematical modelingTool useOpen-ended tasks
FROM THE PAPER / Publication figure · Findings of EMNLP 2025

ModelingBench, ModelingAgent and ModelingJudge

Original ModelingAgent framework figure with contest-data construction, external tools, shared-memory specialist agents, and expert-informed modeling evaluation.View original
Original overview from ModelingAgent: data and task construction, the tool-using multi-agent framework, and ModelingJudge. The website groups the research question at the Agent scale; the actual framework is multi-agent.
Read the paper

Turning a problem into a model

ModelingAgent addresses open-ended problems that require a mathematical formulation, computational tools, and a defensible report. Unlike a single-answer exercise, a practical modeling task can admit several valid solutions with different assumptions and trade-offs.

A connected research framework

ModelingBench supplies competition-derived problems across domains such as transport, ecosystems, and operations. ModelingAgent connects four roles—Idea Proposer, Data Searcher, Modeling Implementor, and Report Writer—through shared memory. A Critic Module supports iterative refinement, while tools provide data access and code execution.

ModelingJudge assesses the resulting reports through multiple expert perspectives. Its evaluation considers completeness, coherence, grounding, and innovation rather than checking only a final number.

Place in the research agenda

This work appears at the Agent scale because it asks how model capability becomes effective task execution. Its actual architecture is multi-agent. The research map organizes questions; the paper describes the system in full.

Paper and contribution

Hongyi Du is a co-author. This brief follows ModelingAgent: Bridging LLMs and Mathematical Modeling for Real-World Challenges, including the framework and evaluation design. The work appeared in Findings of EMNLP 2025. Implementation details are available in the repository.

OPEN CONVERSATION / ModelingAgent

Continue the conversation.

Questions and perspectives on this page are welcome.

Prefer a private conversation?

Loading comments…

Leave a public comment

This conversation belongs to ModelingAgent. Comments appear only after review. For contact details or personal matters, use the private message form.

Public · reviewed before appearing