Model-Based Agentic Software Engineering

How should we engineer software when implementation becomes abundant but engineering judgment remains scarce?

Capable agents have made it much cheaper to produce working code. They have not made it cheaper to decide what should be built, which obligations govern it, how to read the resulting evidence, or whether the system is acceptable to ship.

MAGE studies the engineering structures that make delegated implementation governable: making consequential knowledge explicit as models, giving selected obligations authority over what agents produce, validating realizations against those models, and turning recurring human judgment into structures that later work can reuse.

Research

Scale creates the reasoning problem; commodity intelligence changes its economics. Modeling makes engineering properties explicit, alignment gives selected obligations authority, and governance conversion turns recurring
Scale creates the reasoning problem; commodity intelligence changes its economics. Modeling makes engineering properties explicit, alignment gives selected obligations authority, and governance conversion turns recurring judgment into structures later work can reuse. From the MAGE book.

The MAGE book and course

The book, the course mirror, the detailed framework, and adoption guidance are maintained separately.

Explore MAGE →

Writings

2026
ICSE Journal Ahead Workshop (JAWs)
A registry for agents, treating discoverability, verification, and reproducibility as properties an agent ecosystem has to provide rather than properties individual users establish.
ICSE Journal Ahead Workshop (JAWs)
Extends optimization from local edits to reasoning across a system, which is where delegated work starts to need explicit models rather than context.
Proceedings of the 23rd International Mining Software Repositories Conference -- Mining Challenge track (MSR-MiningChallenge)
Characterizes what agents actually do when asked to optimize code, which is the empirical ground for claims about where judgment still has to sit.
2025
arXiv
Showed that language models can improve the measured performance of real software systems, and established the measurement setup the later agent-optimization work builds on.
In preparation
The book-length statement of the framework.
Software Supply Chains are Dead: Use-Case-Oriented Regeneration
Argues that when regeneration is cheap, reuse decisions change shape. The position that connects the supply-chain work to the agentic setting.

Funding and support

This work has been supported by:

US National Science Foundation
CAREER: PTM-SEER: Software Engineering Foundations for Re-Using Pre-Trained Neural Models (#2541917)
US National Science Foundation
RFE: Research: Developing and Piloting a Prompt Engineering Competency Framework for Software Engineering Education (#2452533)