Configuring Autograders
Pawtograder’s autograder is a GitHub Action,pawtograder/assignment-action,
that runs inside each student repository on every push. It overlays the
student’s submission onto your grader repo, runs the linter, build, and
instructor test suite (and optionally the student’s own tests with
mutation/coverage analysis), then reports per-test results, scores, and
artifacts back to Pawtograder. This page is the reference for the
pawtograder.yml config that drives that flow.
The
pawtograder.yml schema is published at
https://raw.githubusercontent.com/pawtograder/assignment-action/refs/tags/v4/pawtograder.schema.json.
Reference it from the top of your YAML to get IDE autocomplete:On This Page
grade.yml Workflow
The workflow file students run, what counts as a submission, plus action inputs and outputs.
Submission Limits
The 5-per-24-hours cap every assignment starts with, and what counts against it.
pawtograder.yml Configuration
Top-level reference for
build, gradedParts, submissionFiles, and friends.Dependencies
Gate parts or units on prior results.
Feedbot
LLM-generated hints attached to failing tests, including per-test hints from custom graders.
Examples
Working
pawtograder.yml files for Java, Python, and mutation testing.Empty Submission Detection
How Pawtograder rejects submissions that haven’t been changed from the starter.
Submission Viewer
How files and grader artifacts render in the UI.
Rerunning the Autograder
Regrade existing submissions against a chosen grader version.
Test Insights & Bulk Regrading
Find systemic test failures and regrade affected submissions in bulk.
Running the Grader Locally
Iterate on your grader outside GitHub Actions.
Architecture Overview
Advanced: the three repos involved and the action’s step-by-step flow.
Running a Forked Action
Advanced: point students at your fork of the action.
The grade.yml Workflow
The handout repository ships with a .github/workflows/grade.yml that is
cloned into each student repository. You must edit this file to install any
language toolchains or dependencies your build needs before the action runs.
The action itself only does grading — it does not install Java, Python,
Node, etc.
A minimal Java workflow looks like this:
Runner label and action version.
runs-on: ubuntu-latest and
pawtograder/assignment-action@v4 are the portable defaults: they work on any
GitHub-hosted runner and on any Pawtograder deployment. Both v4 and v3 of the action
are live and maintained, so a course template pinned at @v3 is current, not stale.A self-hosting deployment normally substitutes its own self-hosted runner label here. The
Khoury deployment, for example, ships handouts with runs-on: khoury-course-runners, and
offers a resource-class ladder — khoury-course-runners (1 CPU, 1.5 GB RAM),
khoury-course-runners-md (2 CPU, 3 GB), and khoury-course-runners-lg (4 CPU, 6 GB).
Larger classes limit concurrency, so pick the smallest one that can run your build.Follow whatever your own course template pins for both the runner label and the action
version. Do not edit a working template to match the example above.submission/. The action downloads
the grader into a sibling grader/ directory. Both you and the student can
view the run output (including the action’s job summary table) under the
Actions tab of the student repo.
What Counts as a Submission
Theon: block above is what decides when a student submits. It is the workflow’s own configuration, not something Pawtograder imposes, so editing it changes the submission rule for the assignment.
With the push: trigger in place, every push to main is a submission. The grading job runs against the last commit of the push, so a student who pushes three commits at once produces one submission, for the third. Students do not need to write anything special in the commit message.
That trigger is what the handout template ships with. pawtograder/template-assignment-handout declares both workflow_dispatch: and push: branches: [main], and its own comment names the push: block as the switch: “To change the behavior of submitting on every push (vs just when commit has message submit or when the submission is manually requested), change this”. So unless you or a previous instructor edited grade.yml, every push is a submission on your assignment.
Delete the push: block if you would rather students submit deliberately. Pawtograder then dispatches the workflow only when the head commit message of a push contains #submit, matched case-sensitively — so #Submit and #SUBMIT do not count. Warn students if you do this, because nothing in the interface tells them the marker is required. You can also narrow the trigger rather than remove it, by adding branches or restricting paths; the rule is the same either way, in that whatever the on: block matches is what becomes a submission.
Assignments configured without an autograder have no grade.yml at all. There, every push to the default branch is collected as a submission for hand grading, and neither trigger applies. Pushes to other branches are recorded in the commit history but are not submissions. See Configuration → Autograder.
Students Must Not Edit grade.yml
Pawtograder records a hash of the handout’s grade.yml. When a submission arrives, it hashes the copy in the student’s repository and compares the two. On a mismatch the submission is rejected, and the student is told:
Your .github/workflows/grade.yml file does not match the expected contents for this assignment (it may have been edited, or your copy may differ from the handout). Restore the file from your repository’s initial commit, then commit and push.
The same message is recorded as a public error you can see under
Workflow Runs → Grading Errors. The handout template’s own opening comment warns students that tampering could expose the instructor solution and that instructors will be notified.
Three carve-outs: the comparison also accepts a whitespace-stripped hash, so reindenting alone does not trip it; #NOT-GRADED submissions skip the check; and graders and instructors are allowed through with a recorded warning rather than a rejection.
Action Inputs
Action Outputs
Submission Limits
Autograded submissions are rate-limited, and not as an opt-in: every assignment is created with a limit of 5 submissions per rolling 24 hours. The limit is applied automatically to every new assignment, and nothing on the assignment creation form mentions it. An instructor who has never opened this page still has a limit in force.Where to change it
On the Configure Autograder page, not on Edit Assignment:- Maximum number of submissions per student (count) — a whole number, or blank for no limit.
- Maximum number of submissions per student (time period) — the dropdown offers exactly No limit, 10 minutes, 1 hour, 24 hours, and 48 hours.
0: a 0 also disables enforcement, but it reads as a deliberate zero to the next person who opens the page.
What counts against the quota
- The window is a rolling lookback from the moment of the push, not a calendar day.
- A submission whose autograder run has not finished yet does count — the exclusion is only for a finished run that scored zero. Students who resubmit while a run is still queued are spending quota.
#NOT-GRADEDsubmissions do count. There is no exclusion for them.- On a group assignment the quota is shared by the whole group.
- Regrade re-runs and staff-triggered submissions are exempt.
The count students see can disagree with the limit
The remaining-submissions count on the student’s assignment page is calculated separately from the limit that actually blocks a push. The two normally agree, but on group assignments they can drift apart, so treat the student-facing number as approximate when you are fielding a question about it.What the student sees
Two different strings, for the same setting:- The GitHub Actions run fails with
Submission limit reached (max N submissions per 1 day). Please wait until MM/dd/yyyy HH:mm to submit again. - The banner on the assignment page describes the same 86400 seconds as “per 24 hours”, and adds that submissions scoring zero do not count.
The pawtograder.yml Configuration
pawtograder.yml lives at the root of the grader/solution repo. There is
currently exactly one grader type (grader: overlay). The schema requires
three top-level sections, and Pawtograder itself requires a fourth:
grader— the grader type.overlayis the only accepted value.build— how to build, lint, and test the project.submissionFiles— which files from the student repo are collected and overlaid onto the grader.gradedParts— what tests are worth what points, organized into parts. The schema does not list this as required, but the platform does: the Configure Autograder page refuses to load a config without it and reports “Invalid pawtograder.yml structure: Missing required field: gradedParts”, and the in-app config editor blocks saving one. Treat it as required.
feedbot,llm,mutantAdvice— LLM-based features (see Feedbot and Mutation Test Units).maxImplementationHints: N— across all regular units, show full output for at mostNfailing tests. Once the limit is reached, additional failing tests still count against the score but are summarized as “N additional failing tests not shown.” This is a running total across the entire submission, so put the most important parts first ingradedPartsif you care which hints “win.” For per-unit suppression, usehide_output: trueon a regular unit (see below).maxMutantHints: N— capsmutantAdvicehints shown to students; covered with the mutation example in Mutation Test Units.fallbackFiles— provides defaults for files the student didn’t submit (see below).
build
The only required field is preset. The other fields are conditional:
script_info and venv apply to the python-script preset,
student_tests controls mutation/coverage features, and timeouts_seconds
overrides the built-in timeouts.
Presets
java-gradle— Builds with./gradlew test, uses Surefire XML for test results, JaCoCo for coverage, Checkstyle for linting, and Pitest for mutation testing. The grader repo must contain a workingbuild.gradle.python-script— Runs the shell commands you provide inscript_info(see below). Use this when you want full control over how tests, coverage, and mutation are produced.none— Disables building, linting, and testing entirely. The action still records the submission and runs handgrading flows. Useful for write-only / artifact-only assignments.
Linter
policy: ignore— lint errors are reported in the grading summary but tests still run.policy: fail— if the linter finds errors, the rest of grading is skipped and the student receives a zero. A finished run that scored zero does not count against the assignment’s submission limit — and every assignment has one, so this exemption matters.
student_tests
Controls what to do with the student’s own test suite. Tests are run in two
contexts:
Mutation analysis under
instructor_impl only runs if the student’s tests first pass against the instructor’s reference solution. The rationale is that if a student’s tests fail against a known-correct implementation, they’re asserting wrong behavior, so their mutation score isn’t meaningful. The action surfaces those failing tests in a dedicated “your test suite contains incorrect tests” message.timeouts_seconds
All sub-fields are optional; the defaults are:
venv and script_info (Python preset)
For the python-script preset, you supply the shell commands the builder
should run for each phase:
script_info fields are required even if a given phase isn’t used —
provide a no-op command if you don’t need one. cache_key keys the cached
venv across runs; bump it when requirements.txt changes.
artifacts
A list of files or directories the grader will produce and upload to the
submission view. Each entry has a name (shown in the UI), a path
(relative to the grading workspace, or absolute), and optional data (a
free-form object — for example, { "format": "zip", "display": "html_site" }
tells the UI to render a directory as a navigable HTML site).
report_mutation_coverage or
report_branch_coverage is enabled) are added to this list at runtime —
you don’t need to declare them yourself.
gradedParts
name and an array of gradedUnits. Optional fields:
hide_until_released: true— students cannot see this part’s score or test output until the submission is released for grading.dependencies— see Dependencies below.hideFeedbot: true— Feedbot will not generate hints for any failing test in this part.
Regular Test Units
testsmay be a single string or an array of strings. Each string is matched as a prefix against the fully qualified test names emitted by the test runner (for JUnit,package.ClassName.testMethod). A prefix likeCreditCardPublicTest.matches every method on that class.testCountis the number of tests you expect to match. Setting this explicitly is intentional: it prevents a typo in a prefix from silently awarding full marks for zero tests.pointsis the unit’s max score.allow_partial_credit— defaults tofalse. When false, the student earnspointsonly if all matched tests pass and the number of passing tests equalstestCount. When true, the student earnspoints * (passing / testCount).hide_output: true— replaces student-visible test output with “Output for this test is intentionally hidden.” The full output is still recorded ashidden_outputand is visible to staff.hideFeedbot: true— suppresses Feedbot hints for this unit only.
Mutation Test Units
locationsis an array of strings. A class name counts every mutant whose location starts with that class. Any entry that does not match a class is then matched as a substring against each mutant’s name, which underjava-gradleis the Pitest mutator name, soMathMutatorselects every math mutant regardless of class.- Scoring uses either
breakPointsorlinearScoring, not both:breakPoints— array of{ minimumMutantsDetected, pointsToAward }, ordered from the highestminimumMutantsDetectedto the lowest. The unit awards the points of the first entry in list order whose threshold the student met, and it reports the unit’s maximum score as the first entry’spointsToAward.linearScoring: { total_faults, points }— awards(detected / total_faults) * points, rounded to two decimal places.
hideFeedbot: true— same meaning as on regular units.
Line-range locations (
ClassName-startLine-endLine) only match mutants whose reported location has the form Class:startLine:endLine. The java-gradle preset reports Pitest locations as Class:lineNumber, with no end line, so a line range never matches anything under Pitest. Restrict a mutation unit to part of a class only if you supply your own python-script mutation runner that emits three-part locations; otherwise use whole-class locations.mutantAdvice is configured at the top level, mutants the student
didn’t detect can show a personalized hint. Advice is matched by
targetClass: the action splits each surviving mutant’s name on spaces and
compares the last word to the targetClass of every mutantAdvice entry.
maxMutantHints caps the total number of mutantAdvice hints shown
across all mutation units in a single submission. Like
maxImplementationHints, it’s a running total — order gradedParts so
the most important parts come first if you care which hints “win.” Omit
it to show all available hints.
submissionFiles
filesare the source/implementation files that get overlaid onto the grader for the instructor test runs.testFilesare student-written tests; they are kept separate so they can be overlaid only when the action wants to grade the student’s own tests, and so that mutation/coverage analysis has a clean target.- Patterns are GitHub Actions globs —
**for “any subdirectories”,*for “any name in this directory”. You can list a literal file alongside a glob to make that file required.
fallbackFiles
Optional. The path (relative to the grader repo) of a directory whose
contents should be copied into the grading workspace for any file the
student did not submit. Useful when students may delete files that your
test harness expects to exist.
Dependencies
BothgradedParts and gradedUnits accept a dependencies array. If any
dependency is not met, that part (or unit) is replaced in the feedback with
a message explaining which dependency failed instead of the actual grading
output.
A dependency may be written in any of three forms:
- If
minScoreis omitted, the dependency requires the maximum score for the referenced part or unit. minScoreis a raw score, not a percentage.- When a part’s dependencies fail, the entire part is replaced with one feedback entry.
- When a unit’s dependencies fail (but the part’s are satisfied), only that unit is replaced.
Feedbot
Feedbot is optional, LLM-generated feedback that the grading server can attach to failing tests. When enabled inpawtograder.yml, the action
includes an llm block on each failing test result so the grading server
knows which model and account to use.
enabledis required for Feedbot to run at all.provider,model,account, andspec_urlare all required whenenabled: true. If any are missing, Feedbot is disabled for the run and a warning is written to the visible output.spec_urlshould point to a markdown file with the assignment spec. The action fetches it at grading time with a 10-second timeout; if the fetch fails, Feedbot is disabled for that run and the failure is logged.promptselects the response strategy. The two built-ins arechain_of_thought(default) andchecklist. Any other string is used as a free-form custom strategy instruction; the embedded assignment spec and the underlying role/rules are not changed.rate_limitis optional, and so is each field inside it, but the fields do not all default the same way.assignment_totalandclass_totalare lifetime counts of Feedbot responses for the assignment and for the class, with no time window, and each is enforced only if you give it a number: omit one, or set it to0, and that cap does not apply.cooldownis the exception — omit it and you still get a five-minute cooldown, andcooldown: 0is the only way to run with none.accountselects which set of provider credentials the server uses — for example,account: cs2100will look upOPENROUTER_API_KEY_cs2100(falling back toOPENROUTER_API_KEY) when Feedbot dispatches the call.
{account} is the value of the account
field):
openai—OPENAI_API_KEYorOPENAI_API_KEY_{account}.azure—AZURE_OPENAI_ENDPOINTplusAZURE_OPENAI_KEY(orAZURE_OPENAI_KEY_{account}).anthropic—ANTHROPIC_API_KEYorANTHROPIC_API_KEY_{account}.openrouter—OPENROUTER_API_KEYorOPENROUTER_API_KEY_{account}. Use models likeopenai/gpt-4o-mini,anthropic/claude-3-haiku,google/gemini-pro.
hideFeedbot: true on the part or unit.
Per-Test Hints from Custom Graders
When Feedbot is enabled, the action automatically emits anextra_data.llm block on each failing test so the server knows which
model/account to invoke. If you are writing a custom python-script
grader and want to provide per-test hint configuration directly (rather
than going through the feedbot block), you can author the llm block
in your test output yourself:
provider, model, and account fields use the same key lookups
documented above.
Examples
Java with Gradle and JUnit
Python with Custom Scripts
Java with Mutation Testing (Pitest)
This example also grades the student’s own tests for fault-detection strength. The Gradle plugin used isinfo.solidsoft.pitest.
build.gradle enables the Pitest plugin:
Empty Submission Detection
Pawtograder can compare the files it collects from a submission against the starter code in the handout, and where it can, it rejects a submission whose collected files are identical. Two separate things decide what happens: whether the assignment gives the check anything to compare, and how the submission arrived.When the check applies
The comparison covers exactly the files matchingsubmissionFiles in the
autograder configuration, so:
- An assignment that declares
submissionFileshas empty submission detection in force. Nothing in the assignment editor turns that rejection off. - An assignment that declares no
submissionFilesgives the check nothing to compare, so it does not apply and every push is accepted. Repo-only assignments carry no autograder configuration and therefore never see empty submission detection.
submissionFiles, it also respects whatever
scope your autograder configuration defines: an edit confined to a file no
pattern matches does not register as a change.
What a rejection looks like
On an assignment where the check applies, what the student sees depends on how the submission arrived:- Submissions graded by GitHub Actions. The submission is discarded and the run fails with “Empty submissions are not permitted for this assignment. Please commit your changes before submitting.”
- Submissions Pawtograder collects directly from a push. The submission stays in the student’s history, inactive and with no files, and carries a student-visible error explaining that the commit was not recorded because it is identical to the starter code. The assignment page labels that submission “Error”.
Submission Viewer
Submission files and grader-generated artifacts are displayed side by side in the submission viewer. Submitted files:- Text files render with syntax highlighting.
- Markdown files (
.md,.markdown) render as formatted HTML with code-block highlighting, images, tables, and links. - Binary files (images, PDFs, executables) are stored with the submission and exposed as a download button alongside file metadata.
data object on each artifact:
- Plain-text artifacts (
.txt,.log) render with line numbers and syntax highlighting. - Markdown artifacts render as formatted HTML.
- Directory artifacts with
data: { format: zip, display: html_site }are uploaded as a zip and rendered as a navigable HTML site (this is how Jacoco/Pitest HTML reports show up). - Other binary artifacts are exposed as downloads.
annotation_target: artifact on the rubric check and naming the artifact
in the artifact field. See the
Rubrics documentation for details.
Rerunning the Autograder
You can rerun the autograder on an existing submission from the assignment page, the test-insights page, or an individual submission. Reruns keep the original submission record (same timestamp, same submission count) and replace the autograder result. Each rerun lets you choose which grader version to use:- The current grader (latest commit on the grader repo’s default branch).
- A specific commit from the recent history list.
- A manual SHA, for precise version control.
Test Insights and Bulk Regrading
The Test Insights view groups identical test failures across the whole class so you can quickly find systemic problems (a flaky test, an ambiguous spec, an off-by-one in your reference solution). From any error group you can:- See the number of affected submissions and their average score.
- View and copy the email addresses of affected students.
- Pin globally important issues so they remain visible across assignments.
- Launch a regrade with those submissions preselected on the rerun-autograder dialog.
Running the Grader Locally
You can run the grader against a local solution and a local submission without involving GitHub Actions or the Pawtograder server. From a clone ofpawtograder/assignment-action:
pawtograder-grading/
directory in your current working directory; delete it between runs
(or you may hit EACCES errors copying files).
Architecture Overview
When the action runs in the student repo, it:1
Authenticates with the grading server
GitHub issues an OIDC token to the workflow. The action sends that token to the grading server, which verifies it (so it knows which repo and commit the request came from), runs security checks, registers a new submission, and returns a one-time download URL for the matching grader repository tarball.
2
Downloads the grader and reads pawtograder.yml
The action extracts the grader tarball alongside the student’s checkout and reads
pawtograder.yml from the grader repo. The config selects an “overlay” grader and a build preset (java-gradle, python-script, or none).3
Overlays student files onto the grader
For each glob in
submissionFiles.files and submissionFiles.testFiles, the action deletes the matching files in the grader checkout and copies the student’s files in. This is the “overlay”: the grader repo provides the harness, the student’s files are layered on top.4
Lints, builds, and runs instructor tests
The selected builder runs the linter (if configured), then a clean build, then the instructor test suite. Results are parsed into per-test pass/fail records. If
linter.policy: fail is set and the linter fails, or if the build fails, grading stops and a zero is recorded.5
Optionally runs student tests and mutation analysis
If
student_tests is configured, the action resets the grader’s solution files, layers in only the student test files, and runs them against the instructor implementation (and optionally mutation testing). It can also run the student’s tests against the student’s own implementation to report branch and mutation coverage.6
Scores parts and units, resolves dependencies
Scores are computed for every
gradedUnit, summed into gradedPart scores, and then dependency rules are applied — units or parts whose dependencies aren’t satisfied are replaced with a message instead of their actual results.7
Submits feedback and uploads artifacts
The action submits the tests, lint output, and logs back to the grading server. If the grader emitted any
artifacts, they are uploaded to Pawtograder’s file storage via the signed URLs returned by the server. A summary table is also written to the GitHub Actions job summary.handout_notice and
the action exits successfully without grading.
Running a Forked Action
The grading action is fully open source atpawtograder/assignment-action,
so if the pawtograder.yml schema documented above isn’t expressive enough
for what your assignment needs, you can fork the action and point your
assignment’s grading workflow at your fork instead. Common reasons to do
this include adding a new build preset, changing how scores are computed,
or wiring up custom artifact handling.
To use a fork, change the uses: line in grade.yml to point at your
fork and the ref (tag, branch, or commit SHA) you want students to run: