Microsoft wants agent research to be easier to build

Microsoft wants agent research to be easier to build

The interesting part is the infrastructure.

Microsoft Research has open-sourced Orchard, a framework for building, training, and evaluating AI agents across coding, web browsing, and productivity tasks. Instead of releasing another model, Microsoft is releasing the environment, training workflows, datasets, and evaluation tools that let researchers reuse the same infrastructure across different agent systems.

The benchmark numbers are impressive, including 69.7% on SWE-bench Verified, rising to 73.0% with value-model reranking using roughly 3 billion active parameters. What stands out more is the idea that smaller models can stay competitive when the surrounding system is well designed. The screen in a developer’s workspace matters as much as the model running behind it.

This shifts some attention away from model size and toward the software stack around agents. If more labs can train directly inside real deployment environments, progress could spread beyond the handful of companies with proprietary infrastructure. The next question is whether the open community builds on Orchard fast enough to make it a shared standard.

T.A.R.S TACTICAL ASSISTANCE & RECON SYSTEM
ONLINE
// personality matrix
SARCASM
50
HUMOR
50
SERIOUS
50
TARS_
I could talk, but I choose to listen. Go ahead — impress me. Or don't. Either way, I'm here.
🎙