Ends in 7 weeks
114 participants
69 submissions

SWE MERA Challenge

Fresh, uncontaminated real-world software engineering tasks across Go, Python, PHP, and more.

Existing benchmarks for this task face a fundamental dilemma: they are either compromised by data contamination, or exposed to agents that exploit internet access to copy ready-made commits directly from GitHub.

To resolve this, we introduce a two-tier leaderboard guarded by a network allow-list that grants access to standard software repositories, like PyPI, while blocking leakage of solutions. The same design also accommodates self-hosted open-source models, which participants can run on organizer-provided GPUs without any external network access.

Public (Stage I)

The public tier is open for continuous, virtually unlimited debugging, with no restrictions on either the models or the harnesses used.

Private (Stage II, Stage III)

The private tier, by contrast, permits only three scoring attempts per day against a hidden dataset — striking a deliberate balance that rewards genuine model generalization and protects the academic integrity of the final evaluation. Stage II still allows debugging: for example, the allow-list is extended dynamically by organizers as participants request legitimate tools for their agents.

We provide a baseline that shows how to run and submit a solution. You can replace either the model or the harness inside this baseline.

Competition Timeline

PeriodPhaseLanguages
Jul 13 -- Oct 12, 2026Stage I (Public)Go, Python, PHP
Oct 12 -- Nov 15, 2026Stage II (Private)Go, Python, PHP, Language 1
Nov 16 -- Nov 18, 2026Stage III (Private)Go, Python, PHP, Language 1, Language 2
Nov 2026Final private scoring---

Metric

The leaderboard metric is pass@1: the fraction of tasks solved on a single attempt.

  • Each trial is scored by the verifier with a binary reward: 1 if the agent's patch makes the hidden evaluation tests pass, 0 otherwise (written to <trial>/verifier/reward.txt).
  • Since the rules allow exactly one attempt per task (--n-attempts 1), pass@1 is simply the mean reward across tasks: pass@1 = (number of trials with reward 1) / (number of tasks).
  • This is the same value Harbor reports as Mean in the run summary table and in the run-level result.json (stats.evals.<eval>.metrics[0].mean), so you can see your provisional score locally before submitting. The submission checker recomputes it independently from the per-trial verifier/reward.txt files.

Example: a run over 3 tasks with rewards 1, 1, 0 has pass@1 = 2/3 ≈ 0.667.

Requirements

  • uv, Docker, Git, and 60 GB of free disk space (each task ships its own Docker image of 1–2 GB; a full run pulls all of them).
  • On Apple Silicon: task images are linux/amd64. Docker (or Colima) runs them via emulation — make sure Rosetta/qemu is available. Pre-pulling images with --platform linux/amd64 (see Troubleshooting) avoids build-time surprises.

Baseline

More details can be found here: SWE-MERA-CHALLENGE

Install

uv pins Python 3.12, creates the virtual environment, downloads the interpreter when necessary, and installs the pinned dependencies:

git clone https://github.com/MERA-Evaluation/SWE-MERA-CHALLENGE.git
cd SWE-MERA-CHALLENGE

uv python pin 3.12
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
uv run --python .venv/bin/python -- harbor --version

Model configuration

Store provider credentials in .env, then select the model and agent using the harbor run arguments below.

cp configs/.env.example .env
vim .env

Run a single task (recommended first step)

Before a full run, validate your setup end-to-end on one task. This builds one image, runs the agent, and produces a scored trial in a few minutes:

uv run --python .venv/bin/python -- harbor run \
  --path tasks \
  --include-task-name py-009 \
  --agent mini-swe-agent \
  --model openrouter/qwen/qwen3.7-flash \
  --env-file .env \
  --jobs-dir jobs/debug \
  --n-attempts 1 \
  --n-concurrent 1 \
  --agent-setup-timeout-multiplier 10 \
  --yes

The reward for each trial is written to <trial>/verifier/reward.txt (1 = solved, 0 = not). Inspect a finished run with harbor view jobs/debug.

Run the full benchmark

uv run --python .venv/bin/python -- harbor run \
  --path tasks \
  --agent mini-swe-agent \
  --model openrouter/qwen/qwen3.7-flash \
  --env-file .env \
  --jobs-dir jobs/qwen3.7-flash-run-001 \
  --n-attempts 1 \
  --n-concurrent 2 \
  --agent-setup-timeout-multiplier 10 \
  --yes

Estimated cost: Approximately 1.5 USD for a full run with this model.

Submit results

Point the packaging script at the timestamped job directory (the run itself, not its parent):

scripts/create-submission.sh jobs/run-001/2026-08-30__05-45-50

This creates sample_submission.zip. Upload this file using the "Submit solution" button at the top of this page.

What the checker expects. The ZIP must contain exactly one Harbor run: the contents of a single timestamped job directory, with the run metadata files at the archive root and one directory per trial:

sample_submission.zip
├── config.json          # run config
├── lock.json            # environment/task lock
├── result.json          # run-level results and reward stats
├── job.log
└── <task>__<id>/        # one directory per trial
    ├── result.json      # trial results
    ├── agent/           # agent trajectories
    ├── artifacts/       # includes the produced model_patch.diff
    └── verifier/
        ├── reward.txt   # 1 or 0 — the trial score
        └── report.json, parse_result.json, ...

Common mistakes that lead to rejection:

  • Zipping the parent jobs/ directory or several runs at once — the checker requires exactly one run (harbor_runs_found: 0 or more than 1).
  • Packaging hand-picked files (e.g. a CSV of patch files) instead of a genuine Harbor job directory — the checker reads the run's result.json and per-trial verifier/reward.txt, nothing else.
  • Re-packaging a run after manually editing results — keep the original job directory intact; organizers may request it for verification (see Rules).

Troubleshooting

Problems we hit ourselves while validating the baseline:

  • Agent setup times out (default 360 s). Installing agent dependencies inside the container can be slow on cold caches or emulated CPUs. Pass --agent-setup-timeout-multiplier 10.
  • Image pulls fail intermittently during the run. The registry (auth.docker.io) sometimes times out mid-build. Pre-pull the task images before starting:
    docker pull --platform linux/amd64 rewive/swe-mera-challenge:pallets-click-3208
    Image tags follow the pattern rewive/swe-mera-challenge:<owner>-<repo>-<pr-number> and are listed in each task's YAML file under tasks/.
  • Apple Silicon (arm64). All task images are linux/amd64; use Docker Desktop or Colima with emulation. Everything (build, agent run, verifier) works under emulation, just slower — budget roughly 2–3× the wall-clock time.
  • Checking where a run went wrong. Look at <trial>/trial.log (agent/verifier lifecycle), <trial>/verifier/command_test.log (test output), and <trial>/verifier/reward.txt (final score). harbor view <jobs-dir> gives a summary table.

Rules

  • Give each task exactly one attempt (--n-attempts 1).
  • Submit only patches produced by the agent during the run.
  • Do not use gold patches, evaluation tests, or test-specific workarounds.
  • Retain the Harbor job directory and trajectories for verification.

Support

If you want to ask any questions, there is an official support channel in Mattermost: https://mm.ods.ai/ods/channels/swe-mera-challenge. To access Mattermost use your ODS.ai account (choose authorize via ODS.ai on the Mattermost welcome page).

Our website uses cookies, including web analytics services. By using the website, you consent to the processing of personal data using cookies. You can find out more about the processing of personal data in the Privacy policy