Fresh, uncontaminated real-world software engineering tasks across Go, Python, PHP, and more.
Existing benchmarks for this task face a fundamental dilemma: they are either compromised by data contamination, or exposed to agents that exploit internet access to copy ready-made commits directly from GitHub.
To resolve this, we introduce a two-tier leaderboard guarded by a network allow-list that grants access to standard software repositories, like PyPI, while blocking leakage of solutions. The same design also accommodates self-hosted open-source models, which participants can run on organizer-provided GPUs without any external network access.
The public tier is open for continuous, virtually unlimited debugging, with no restrictions on either the models or the harnesses used.
The private tier, by contrast, permits only three scoring attempts per day against a hidden dataset — striking a deliberate balance that rewards genuine model generalization and protects the academic integrity of the final evaluation. Stage II still allows debugging: for example, the allow-list is extended dynamically by organizers as participants request legitimate tools for their agents.
We provide a baseline that shows how to run and submit a solution. You can replace either the model or the harness inside this baseline.
Competition Timeline
| Period | Phase | Languages |
|---|---|---|
| Jul 13 -- Oct 12, 2026 | Stage I (Public) | Go, Python, PHP |
| Oct 12 -- Nov 15, 2026 | Stage II (Private) | Go, Python, PHP, Language 1 |
| Nov 16 -- Nov 18, 2026 | Stage III (Private) | Go, Python, PHP, Language 1, Language 2 |
| Nov 2026 | Final private scoring | --- |
The leaderboard metric is pass@1: the fraction of tasks solved on a single attempt.
1 if the agent's patch makes the hidden evaluation tests pass, 0 otherwise (written to <trial>/verifier/reward.txt).--n-attempts 1), pass@1 is simply the mean reward across tasks: pass@1 = (number of trials with reward 1) / (number of tasks).result.json (stats.evals.<eval>.metrics[0].mean), so you can see your provisional score locally before submitting. The submission checker recomputes it independently from the per-trial verifier/reward.txt files.Example: a run over 3 tasks with rewards 1, 1, 0 has pass@1 = 2/3 ≈ 0.667.
linux/amd64. Docker (or Colima) runs them via emulation — make sure Rosetta/qemu is available. Pre-pulling images with --platform linux/amd64 (see Troubleshooting) avoids build-time surprises.More details can be found here: SWE-MERA-CHALLENGE
uv pins Python 3.12, creates the virtual environment, downloads the interpreter when necessary, and installs the pinned dependencies:
git clone https://github.com/MERA-Evaluation/SWE-MERA-CHALLENGE.git cd SWE-MERA-CHALLENGE uv python pin 3.12 uv venv --python 3.12 .venv uv pip install --python .venv/bin/python -r requirements.txt uv run --python .venv/bin/python -- harbor --version
Store provider credentials in .env, then select the model and agent using the harbor run arguments below.
cp configs/.env.example .env vim .env
Before a full run, validate your setup end-to-end on one task. This builds one image, runs the agent, and produces a scored trial in a few minutes:
uv run --python .venv/bin/python -- harbor run \ --path tasks \ --include-task-name py-009 \ --agent mini-swe-agent \ --model openrouter/qwen/qwen3.7-flash \ --env-file .env \ --jobs-dir jobs/debug \ --n-attempts 1 \ --n-concurrent 1 \ --agent-setup-timeout-multiplier 10 \ --yes
The reward for each trial is written to <trial>/verifier/reward.txt (1 = solved, 0 = not). Inspect a finished run with harbor view jobs/debug.
uv run --python .venv/bin/python -- harbor run \ --path tasks \ --agent mini-swe-agent \ --model openrouter/qwen/qwen3.7-flash \ --env-file .env \ --jobs-dir jobs/qwen3.7-flash-run-001 \ --n-attempts 1 \ --n-concurrent 2 \ --agent-setup-timeout-multiplier 10 \ --yes
Estimated cost: Approximately 1.5 USD for a full run with this model.
Point the packaging script at the timestamped job directory (the run itself, not its parent):
scripts/create-submission.sh jobs/run-001/2026-08-30__05-45-50
This creates sample_submission.zip. Upload this file using the "Submit solution" button at the top of this page.
What the checker expects. The ZIP must contain exactly one Harbor run: the contents of a single timestamped job directory, with the run metadata files at the archive root and one directory per trial:
sample_submission.zip
├── config.json # run config
├── lock.json # environment/task lock
├── result.json # run-level results and reward stats
├── job.log
└── <task>__<id>/ # one directory per trial
├── result.json # trial results
├── agent/ # agent trajectories
├── artifacts/ # includes the produced model_patch.diff
└── verifier/
├── reward.txt # 1 or 0 — the trial score
└── report.json, parse_result.json, ...
Common mistakes that lead to rejection:
jobs/ directory or several runs at once — the checker requires exactly one run (harbor_runs_found: 0 or more than 1).result.json and per-trial verifier/reward.txt, nothing else.Problems we hit ourselves while validating the baseline:
--agent-setup-timeout-multiplier 10.Image tags follow the patterndocker pull --platform linux/amd64 rewive/swe-mera-challenge:pallets-click-3208
rewive/swe-mera-challenge:<owner>-<repo>-<pr-number> and are listed in each task's YAML file under tasks/.linux/amd64; use Docker Desktop or Colima with emulation. Everything (build, agent run, verifier) works under emulation, just slower — budget roughly 2–3× the wall-clock time.<trial>/trial.log (agent/verifier lifecycle), <trial>/verifier/command_test.log (test output), and <trial>/verifier/reward.txt (final score). harbor view <jobs-dir> gives a summary table.--n-attempts 1).If you want to ask any questions, there is an official support channel in Mattermost: https://mm.ods.ai/ods/channels/swe-mera-challenge. To access Mattermost use your ODS.ai account (choose authorize via ODS.ai on the Mattermost welcome page).
Our website uses cookies, including web analytics services. By using the website, you consent to the processing of personal data using cookies. You can find out more about the processing of personal data in the Privacy policy