Out Mon Oct 5 · due Tue Oct 20, 11:59pm · all five assignments
⚠️ THE LEADERBOARD DOES NOT ACCEPT LATE DAYS AND WILL CLOSE ON TUESDAY, OCTOBER 20, AT 11:59 P.M.
If you have not submitted to the leaderboard by the time it closes, you get 0 points for the leaderboard. Late days apply only to the repo and the video.
The main agent runs in the agent container, built from the course's dispatcher/agent.Dockerfile, with your src/ mounted in. The main agent loop starts in this container at the very beginning, and memory, subagents and any other features should live in it as well. The starter_code folder is mounted read-only at /madsOpt, and your agent can write to only two places in it: .tasks/ (its working directory) and madsOpt_logs/ (where logs go). Anything else it needs to write (scratch files, caches) goes elsewhere inside the container, e.g. /tmp.
Every task runs in its own task container, built from that task's own image: the repository checked out at the task's starting commit, with the toolchain and test dependencies the task needs. This container holds the code your agent reads and modifies. Some Go tasks need modules that the fix must add. Their task containers come with a pre-warmed module cache and have GOPROXY=off set, so go get <module>@<version> and go mod tidy work offline. To see what is available, run ls $(go env GOMODCACHE)/cache/download/<module>/@v/. evaluation_scripts/prepare_images.sh builds the images so that tasks can be solved offline.
Your agent interacts with a facility dispatcher to start a task, generate a patch and evaluate it. Neither container can reach the internet: the agent container's only way out is an egress proxy to the course API, which CS2680_BASE_URL points to, and task containers have none.
The figure below shows where each piece runs and how they talk to each other.
When your agent asks for a task, the dispatcher starts the task container, waits until it is ready, and hands your agent a dict with the task spec and the details for reaching the container:
problem_statement, requirements, interface # the issue and its specification
task # the task index k
sandbox_url # http://<container>:8000
workdir # the repository path inside it, e.g. /app
A tiny HTTP server (src/sandbox_server.py) runs inside the task container with four routes: /health, /exec (run a shell command in the repository and get stdout, stderr and the exit code), /read (a file), /write (a file). Your agent talks to it with plain HTTP POSTs. src/sandbox.py wraps them as Sandbox(url, workdir) with exec, read_text and write_text; everything your agent does to a repository goes through those calls. You can add routes for any features you need (a new tool, a subagent, ...). Do not touch the repo with local open() or subprocess: those act on the agent container, which has no repo.
When your agent believes a task is solved, it produces a patch and asks the facility to grade it:
dispatcher.extract_patch(k) # `git diff` of the task container -> the task's patch folder
# (or dispatcher.submit_patch(k, text) to write your own)
ev = dispatcher.evaluate(k) # -> {"tests_failed": F, "tests_total": T, "attempt": a}
The facility copies the patch onto a fresh copy of the task image (with the same pre-warmed Go module cache, and no network), applies the task's hidden tests, runs them, and sends the agent the total number of tests and how many failed. Your agent then chooses between:
dispatcher.continue_task(k) # keep working: the task container and your edits stay exactly as
# they are, the submitted patch is deleted, and you may call
# evaluate(k) again when you have a new one
dispatcher.done(k, reason, iterations)
# the task is final. Its last evaluation is its result; a patch you
# submitted but never evaluated is graded now; no patch counts as a
# failure. The task container and the patch folder are removed.
next_task() raises DispatcherError when 5 are open, and returns None when no tasks are left.continue_task(k), evaluate(k) and done(k) are refused while a grading of k is still running.DispatcherShutdown, every task still open is finished as if done(k) had been called at that moment (its last evaluation counts; a submitted but unevaluated patch is graded; no patch is a failure), and tasks never opened count as failures. Gradings already running still finish.class Agent:
def __init__(self, dispatcher): ...
def run(self): ...
The runner constructs Agent(dispatcher) and calls run(). After that, you decide the flow: how many tasks to keep open, whether to work on them in parallel, when to evaluate, whether to spend another attempt on a failing task or move on, and what to carry from one task to the next. The starter Agent in src/agentic_loop.py only demonstrates these calls. It takes the tasks one at a time and makes no attempt to fix them, so every task fails until you write your own.
Log your model calls and tool calls through dispatcher.logger(k). The leaderboard reads your turn counts from that trace.
The starter code is in the course assignments repository, under assignment3/starter_code/. Copy that entire folder into the same path in your shared private course repository, including init.sh and package_claude_sessions.py.
Before starting work with Claude Code, run init.sh:
# From your shared course repository:
bash assignment3/starter_code/init.sh
This copies package_claude_sessions.py to the repository root and creates CLAUDE.local.md there with a workflow for all work in the repository. Claude Code commits your changes together with the repository's session logs in claude-sessions.tar.gz at the repository root.
After initialization, both helper files live at the repository root:
course-repository/
├── CLAUDE.local.md repository workflow instructions for Claude Code
├── package_claude_sessions.py session packager copied by init.sh
└── assignment3/
└── starter_code/ starter files listed below
The shared workflow runs python3 package_claude_sessions.py --path . --output claude-sessions.tar.gz from the repository root to include sessions from every assignment and other repository work. For the Assignment 3 submission, run python3 package_claude_sessions.py from the repository root to create assignment3/a3-sessions.tar.gz. The remaining instructions use paths relative to assignment3/starter_code/.
You can modify anything under src/ and add features there. The rest (dispatcher/, madsOpt.py and evaluation_scripts/) is the same as what the leaderboard uses. The one exception is dispatcher/student_name.json, where you fill in your name (Step 2); on the leaderboard, the grader's own session id is used instead.
assignment3/starter_code/
├── init.sh run first: copies the session packager and creates CLAUDE.local.md at the repository root
├── package_claude_sessions.py packager source, copied to the repository root by init.sh
├── src/ YOURS: everything here may change (the leaderboard takes only src/)
│ ├── agentic_loop.py class Agent (your entry point): a starter showing every dispatcher, sandbox, model and trace call
│ ├── config.py your harness settings (the course API settings are environment variables, see Step 2)
│ ├── sandbox.py Sandbox(url, workdir): exec / read_text / write_text over HTTP
│ └── sandbox_server.py the HTTP server inside every task container (run_all.py needs its GET /health)
├── dispatcher/ course infrastructure
│ ├── dispatcher.py Dispatcher, the client your Agent calls: next_task, extract_patch / submit_patch, evaluate, continue_task, done, logger
│ ├── egress_proxy.py the reverse proxy to the course API: the agent container's only way out
│ ├── session.py the run's session id, sent with every course API request
│ ├── student_name.json your first and last name, filled in once (see Step 2)
│ └── agent.Dockerfile the agent container image (python + openai, built offline)
├── madsOpt.py course infrastructure: builds the client and calls Agent(...).run()
└── evaluation_scripts/
├── prepare_images.sh one-time setup (needs internet)
├── run_all.py runs your agent on all tasks and grades them
├── evaluate_one.sh grades one patch on a pristine task container
├── make_predictions.py collects the final patches into predictions.json
├── trace_logger.py the JSONL trace your agent writes through dispatcher.logger(k)
├── agent_task_input.json the tasks
└── task_test.json the tests each task is graded with
Step 0. Prerequisites. Linux x86_64; Docker; git; python3 (3.10+) with pip; about 15 GB of free disk for the images. If you find difficulties on your local machine, you can use a CloudLab machine or an AWS machine.
# only if Docker is not installed:
curl -fsSL https://get.docker.com | sudo sh
sudo usermod -aG docker $USER
newgrp docker # or log out and back in
docker run hello-world # check that it works
# only if python3 has no pip ("No module named pip"):
sudo apt update
sudo apt install python3-pip
Step 1. Set up the images.
# From your shared course repository:
cd assignment3/starter_code
bash evaluation_scripts/prepare_images.sh
The script fetches the task data (SWE-bench Pro, at a fixed commit), installs a portable Python for the task containers, downloads the openai wheels, pulls every task's image, builds the agent image, and builds a pre-warmed image for each Go task that requires additional packages. A successful build ends with == all images present.
Step 2. Set the env variables and test it.
export CS2680_API_KEY=...
export CS2680_MODEL_EXPERT=expert # Expert: most capable model; per 1M tokens: $2.00 input, $0.10 cached input, $10.00 output
export CS2680_MODEL_STANDARD=standard # Standard: per 1M tokens: $0.08 input, $0.004 cached input, $0.28 output
export CS2680_MODEL_STARTER=starter # Starter: cheapest model; per 1M tokens: $0.01 input, $0.0005 cached input, $0.04 output
python3 evaluation_scripts/run_all.py --limit 0
run_all.py passes every CS2680_* variable into the agent container, and sets CS2680_BASE_URL to the egress proxy, the agent's only way to the course API. Your agent reads them with os.environ; never hard-code the API URL. It may use any of the three models in one run, and as many of them as it likes.
This command starts the agent container and the proxy, checks the network rules (course API reachable, everything else blocked), starts and health-checks every task container, and then stops without running your agent. Look for EGRESS OK, and for preflight k: ... ok for every task.
The first time you run run_all.py in a terminal, it asks for your first and last name (in Latin letters) and saves them in dispatcher/student_name.json. Every course API request of a run then carries the session id <First>_<Last>_<UTC start time>, e.g. Jane_Doe_20261006T172408Z, which the egress proxy adds; the run prints it and keeps it in run_logs/session_id. If you run without a terminal (e.g. with nohup), fill in that file first, or the run exits before it starts anything.
Step 3. Run it.
python3 evaluation_scripts/run_all.py
Progress is written to run_logs/sequence.log. When the run ends, starter_code contains run_all_results.md (e.g. 3/5 passed), pro_eval/ (the grader's output per task), model_patch_<k>.diff, and madsOpt_logs/<k>/run.jsonl (your agent's traces). The starter agent is only a skeleton and fixes nothing, so expect 0/5 passed from it. Before the next run, archive the previous run's outputs in another folder.
The leaderboard grades your src/ on the evaluation task set, using the same dispatcher/, madsOpt.py and evaluation_scripts/ as the starter code. It takes only your src/ from what you upload. The leaderboard is at https://leaderboard.cs2680.com/
Step 1. Log in and change your password. On the leaderboard page, log in with the Harvard email address associated with your Canvas account. Your initial password is Cs2680-<your 8-digit student ID>; these are the same initial login credentials as for your API account. After logging in, click your name and select Change Password from the drop-down menu.
Step 2. The leaderboard. After you log in, the Leaderboard page shows the top 10 students, each represented by their best submission, and below them your own best submission with its rank. The leaderboard updates as soon as a grading finishes.
Step 3. Submit. On the Submit page, upload a zip of your src/ folder and choose the largest number of tasks this run may grade. The run takes the first k tasks of the evaluation set, or all of them if you leave the field empty.
Step 4. How a submission is measured. Every submission is evaluated on three metrics:
Step 5. Past submissions, cancelling, and your daily budget. The Past submissions page lists your ongoing and past submissions. You can have at most one submission queued or being graded at a time, and that one has a Cancel button. Every student has $30 per day for leaderboard grading, with a $10 limit per submission. If grading reaches either limit, it stops, and whatever it completed up to that point is recorded as the submission's result.
Step 6. Ranking: a skyline. Each student is ranked by their best submission.
Reached the budget, with whatever it had done by then.The submission has three parts: the code (your src/ on the leaderboard), a write-up, and a video.
Zip assignment3/starter_code/src/ and upload src.zip on the leaderboard's Submit page (Part 2, Step 3):
# Run from assignment3/starter_code/:
zip -r src.zip src
Also archive your Claude Code session files as described in What to submit with each assignment and commit them to assignment3/ in your repository. Check that your .gitignore does not exclude the archive. As on every assignment, the session files carry no points of their own. They are the record behind the claims in your write-up.
A PDF document that answers these questions:
The write-up must be entirely your own work, with no AI-generated text. Commit it to your repository as assignment3/writeup.pdf. Your repository's assignment3/ should then contain:
assignment3/
├── starter_code/ your agent code and local run outputs
├── a3-sessions.tar.gz your Claude Code session files
└── writeup.pdf the write-up
In the video, present the same lessons with your face visible and in your own voice. Open your code and move the cursor through it, and navigate your harness: show where each component from the write-up lives, along with the pitfalls and the helpful techniques. Submit the video on Canvas.
| Part | Points | What is graded |
|---|---|---|
| Leaderboard | 80 | The rank of your best submission on the leaderboard (Part 2, Step 6), scored as below. |
| Write-up | 10 | The lessons you learned from the assignment, clearly explained in one page. |
| Video | 10 | A clear 3-minute presentation of the same lessons, showing the code where a lesson concerns your harness. |
Leaderboard score. If r is the rank of your best submission:
score = max(81 − r, 50) if any of your submissions achieves the baseline
score = 81 − r otherwise
The baseline solves 16 tasks in 4 hours within a limit of $10. If any of your submissions achieves the baseline, you get credit for it, even if that submission is not your best. For example, rank 1 scores 80; rank 12 scores 69 whether or not it beats the baseline; rank 40 scores 50 if it beats the baseline and 41 if it does not.
If you find a security vulnerability and report it, you can earn up to 3 bonus points toward your final course score (not just this assignment). Abusing a vulnerability forfeits your submission.