Your code runs on our servers.
Here's exactly what happens.
Every puzzle submission builds and runs server-side with independent safety layers around it, and the grader measures it next to the code, where the browser cannot fake the result. The Labs beta runs on the same foundation.
The pipeline
The journey of a submission
From your keystroke to a graded report, five stops, each one assuming the previous could be wrong.
Write
C# in a full editor, right in your browser. Run the samples as often as you like.
Screen
Every submission is screened before it runs. Code that reaches for known-dangerous APIs is rejected outright.
Compile
Built server-side with the same compiler you use locally, full diagnostics on errors.
Run, sealed
Executed in a disposable, fully isolated sandbox. One per submission.
Measure & grade
Time, memory, and correctness, measured on the server, where they can't be faked.
Seconds, end to end: the snippet goes out, the graded report comes back, per-test timings, budgets, memory, and diffs.
Containment
Inside the sandbox
We treat every submission as hostile, including yours. That's a feature: the same walls that contain malicious code make the grading fair and the timings clean.
$ submit Solution.cs
▸ sandbox spawned fresh · single-use
# security hardening, applied to every run
✓ network isolated
✓ file system read-only
✓ cpu · memory · processes capped
▸ tests complete 148 ms
▸ sandbox destroyed nothing persists
▮
This exact lifecycle runs for every submission, spawn, seal, run, measure, destroy. There is no long-lived server your code shares.
The grading
How each track grades you
Track 01
Algorithms
Correct gets you halfway. The hidden suite scales the input until complexity decides the outcome.
- · Visible sample cases plus a hidden suite with inputs large enough that complexity decides the outcome.
- · Per-test time budgets: correct-but-slow times out where the efficient solution passes with room to spare.
- · Allocation tracking on every run, and on some puzzles a hard allocation budget, where the copy-and-reverse approach fails and the in-place one passes.
The N+1 starter
… ×41 round-trips
✗ Timeout over budget
The set-based rewrite
1 round-trip
✓ Passed 12 ms
Track 02
Database / EF
Your LINQ runs against a real database, and the grader shows you what it really cost.
- · Enough data that inefficiency shows up in the timings, not just in code review.
- · Expand any test to see the queries your code actually produced, and what they cost.
- · Plan-graded puzzles capture the execution plan the engine chose and grade it: full-table reads flagged in red, index usage in green. The right rows the wrong way fails.
method length
limit 25 lines
cyclomatic complexity
limit 8
nesting depth
limit 3 levels
✓ 9 / 9 tests still passing (behavior never changed)
Track 03
Refactoring
Working code you'd hate to inherit. Make it clean, without changing what it does.
- · The behavioral tests pass before you touch anything, and they must still pass when you're done.
- · Structural gates measured from your source: method length, cyclomatic complexity, nesting depth, duplicate blocks.
- · Every flavor of real-world mess (tangled conditionals, god methods, copy-paste blocks, arrow code), including the famous Gilded Rose kata.
Infrastructure
database, email, HTTP
Application
use cases
Domain
pure, depends on nothing
Track 04
Architecture
Multi-file refactoring katas, graded on behavior and design together.
- · Dependency direction, layering, and abstraction boundaries are verified automatically.
- · A wrong dependency fails the submission the same way a failing test does.
- · Feedback in seconds, not in a code review three weeks later.
The adversarial suite
../../etc/passwd
path traversal
✓ denied
evil-example.com
suffix confusion
✓ rejected
http://169.254.169.254/
SSRF
✓ blocked
ada\n[INFO] admin granted
log forging
✓ escaped
Track 05
Secure Coding
Functionally correct, quietly exploitable. Close the hole without breaking the feature.
- · Two suites grade every submission: functional tests prove the feature still works, adversarial tests throw real attack payloads at it.
- · The payloads are the classics that hit production systems: path traversal, SQL injection, SSRF, open redirect, log forging, mass assignment.
- · The catch: the naive fix (blacklist a bad string, strip a tag) is exactly what the adversarial cases are built to defeat. Only a real fix, encode the output, allowlist the input, parameterize the query, passes.
Your refactor
IPricingStrategy.cs
PricingCalculator.cs
CompositionRoot.cs
Working code is supplied. You change the design without changing the behavior.
Behavior suite
The feature still calculates every price correctly
Executable design rules
depends on abstraction
PricingCalculator → IPricingStrategy
no concrete dependency
PricingCalculator ↛ PremiumPricing
sealed policy owner
PricingCalculator is sealed
Track 06
Design Principles
Short katas on the ideas behind the patterns. Keep the behavior, change the design.
- · Working code is supplied, and the behavioral tests must still pass when you're done.
- · Structural rules are checked on your source: immutable types, readonly fields, sealed classes, required abstractions, and which types may reference which.
- · A half fix fails: a private setter still leaves the backing field writable, and the immutability rule catches it.
Your submission
The implementation is supplied. Your inputs and assertions are the solution.
Correct implementation
Every test must pass
Five planted faulty versions
CAUGHT
boundary moved
CAUGHT
branch removed
CAUGHT
operator swapped
CAUGHT
case ignored
SURVIVED
edge survives
4 of 5 caught → grade passed
Track 07
Test Writing
The model inverted: you write the tests, and the grader checks whether they would catch a bug.
- · You get a correct implementation. Your suite has to pass against it first.
- · Then it runs against mutants, copies of the code with one planted bug each, and must fail on at least the number the kata requires.
- · A test that asserts nothing catches no mutants and earns nothing. The grade measures what coverage cannot: whether your tests detect a regression.
A different submission path
System Design grades the topology
The Studio does not compile a diagram or pretend to deploy it. It sends bounded topology data to the trusted grader, which validates the component vocabulary and evaluates authored graph rules with the same result every time.
- Required paths, forbidden bypasses, fan-out, redundancy, and containment are gradeable.
- Request, async, storage, secrets, telemetry, and replication connections keep their meaning.
- A passing Submit opens the production tradeoffs and signals the diagram still cannot prove.
Labs (preview)
From course to check
The Labs beta follows one system through 4 guided lessons. Each lesson changes the live workspace, and executable checks prove the behavior before you continue.
Product
Katabench Labs
The guided, executable practice area.
Course
Transactional Outbox
The beta course takes one small .NET service from a lossy dual write to a reliable outbox in 4 lessons.
Lab
Fix the order service
One service on real PostgreSQL and RabbitMQ, and a focused goal at every step.
Lesson
Make the change
Guided work that advances the lab.
Workspace
Build and run
A disposable environment for your code and live output.
Check
Prove the behavior
Executable verification before the next lesson unlocks.
Your work stays yours
Code and topology submissions are stored so you can see your own history and progress, that's it. We do not publish them, and every code sandbox is gone seconds after it runs. The details live in the Privacy Policy.
See the grading for yourself
One puzzle is all it takes to understand why evidence beats a green checkmark on its own.
Try it free, no signup