News

MIT and Sakana AI's SIFT cuts coding-agent evaluation costs

MIT and Sakana AI's SIFT adds a low-cost judge before full benchmarks, helping prioritize coding-agent changes without replacing deeper testing.
Oct 4, 2026路3 min read
MIT and Sakana AI's SIFT cuts coding-agent evaluation costs

#

Key takeaways

  • SIFT is a research framework from MIT and Sakana AI.
  • It adds a lower-cost judgment step before broad benchmark testing.
  • The system still uses benchmark tests after ranking candidates.
  • Reported results come from specific experimental setups.
  • The article claims no deployment or Morocco-specific use.

What SIFT is trying to solve

VentureBeat reported on October 2, 2026, that MIT and Sakana AI introduced SIFT. The framework targets a practical problem in improving coding agents: evaluation can be expensive. A proposed patch may look promising, but teams still need to check whether it actually improves performance.

The source describes a loop where an agent proposes a patch, a revised version runs coding tasks, and the results guide later changes. SIFT adds an earlier judgment step before broad benchmark testing. That step is meant to reduce wasted evaluation work.

How the framework works

After a modification, the new agent is first run on four coding tasks. This is meant to catch obvious breakage early. The source presents this as a quick filter, not a final verdict.

Next, a language-model judge compares the candidate's code with up to ten strong agents from an archive. The judge does this without seeing benchmark tasks or outcomes. A Bradley-Terry ranking then combines those pairwise preferences.

That ranking helps prioritize more expensive evaluations. It does not replace them. The system can also generate patches, judge them, and run benchmarks in parallel.

What the reported experiments showed

VentureBeat cites experiments on Polyglot, TerminalBench 2.1, and a subset of SWE-bench Verified. In one reported Polyglot setting using o3-mini, SIFT reached 35.1 percent accuracy. The Darwin Godel Machine baseline reached 30.7 percent in the same report.

The article also says a SIFT run without the judge reached 29.8 percent. That suggests the judge can matter in the reported setup. Still, these are benchmark results in specific experimental conditions.

The source also reports that testing a candidate on fifty Polyglot tasks costs about six dollars and 2.6 CPU hours in the paper's breakdown. That figure helps frame the evaluation burden the framework is trying to manage.

Why the cost step matters

SIFT's main idea is not to skip evaluation. It is to make evaluation cheaper and better ordered. The judge gives an early signal that can help decide which candidates deserve deeper testing.

This matters because broad benchmark runs can consume time and compute. A lower-cost ranking step can reduce unnecessary work. The source does not claim that this approach always improves every coding environment.

Limits and governance considerations

The article is careful about scope. The reported numbers come from benchmark experiments, not from a claim of universal improvement. The judge also does not see benchmark tasks or outcomes, which keeps the ranking step separate from the final tests.

That separation is important. It reduces the chance that the judge simply mirrors benchmark answers. It also means the ranking is only a prioritization tool, not a substitute for real evaluation.

Morocco relevance

The source reports no Morocco-specific use. A conditional global lesson is that teams anywhere should treat low-cost ranking as a filter, not a replacement for benchmark testing.

Bottom line

SIFT is a research framework for coding agents that tries to cut evaluation costs. It does this by adding an LLM-based judgment step before broader benchmark runs. The reported results are promising in the cited setups, but they remain experimental.

For readers tracking agent evaluation, the key point is simple. SIFT aims to spend expensive testing only on the most promising candidates. That makes the evaluation pipeline more selective, while keeping final benchmarks in place.

Follow us on Google

Add Intelligence Artificielle Maroc as a preferred source to see more of our relevant stories in Google Search.

Add us as a preferred source
AI platform development

What would you like to build?

We build custom AI platforms, SaaS products, intelligent business applications, and automation systems.

This form is for project inquiries, not general questions about artificial intelligence.

Name *
Work email *
Organization (optional)
Solution *
Short project description *

Related Articles

featured
J
Jawad
路Oct 4, 2026

Secure Web Search in Claude Desktop with Amazon Bedrock AgentCore

featured
J
Jawad
路Oct 4, 2026

Muse Gadgets: Open source hardware for your Muse

featured
J
Jawad
路Oct 4, 2026

NVIDIA DGX Spark 64GB Expands Local AI Options

featured
J
Jawad
路Oct 4, 2026

Sweep thousands of leases with Amazon Quick and Adjudicated Query