MastriaDocs

Lab authoring

Labs are part of your lesson markdown: a fenced ```lab block containing YAML. Because a lab is lesson content, it inherits everything lessons already have — draft vs. published, version history, repo sync. There is no separate lab tool to keep in step.

Learner code runs in the browser — SQL on a real DuckDB engine, Python on a real CPython runtime. Run executes and shows output; Check my work re-runs the code and evaluates your checks against the actual result of the run. A lab you already host elsewhere can be embedded instead, and report results back.

A complete lab

Markdown
```lab
id: filtering-orders          # stable id — progress references it forever
title: 'Lab: Filter orders with WHERE'
estimatedMinutes: 10          # optional, shown in the lab header
language: sql                 # sql (DuckDB) or python (Pyodide);
                              # anything else = code-only checks
setup: |                      # environment script — re-runs before EVERY
  CREATE OR REPLACE TABLE orders (   # learner run, so keep it idempotent
    order_id INTEGER, amount DOUBLE
  );
  INSERT INTO orders VALUES (1, 120.0), (2, 250.0);
starterCode: |                # optional, seeds the editor
  SELECT *
  FROM orders;
exercises:
  - id: add-where             # stable id, unique within the lab
    title: Filter rows with WHERE
    brief: >                  # optional scene-setting, markdown
      Narrow the orders table down to the orders over €100.
    steps:                    # optional numbered instructions, markdown
      - md: Add a `WHERE` clause to the starter query.
        note: WHERE filters rows before SELECT picks columns.
    checks:                   # at least one; ALL must pass
      - type: code_matches
        pattern: '\bWHERE\b'
        flags: i
        fail: "There's no WHERE clause yet — filtering always starts there."
      - type: sql_result_equals
        rows: [[2, 250.0], [1, 120.0]]
        fail: "The query runs, but it's not returning only the orders over 100."
    hints:                    # first hint is free; each failed attempt unlocks one more
      - 'The clause goes after FROM: `SELECT … FROM orders WHERE <condition>`.'
    solution: |               # powers "Apply the solution"; always write one
      SELECT * FROM orders WHERE amount > 100;
```

Exercises unlock one after another; the lab is complete when every exercise has passed.

Tip

You don't have to type this from scratch. In the editor, Add lab stamps a valid skeleton for SQL, Python, code-only, or an embedded lab, and the Exercise, Hint and Check buttons grow it at your cursor.

The six check types

Checks run when the learner presses Check my work. All checks of the exercise must pass. The learner sees them as a checklist, one row per check, passing or failing. A failing check shows exactly one thing beside its row: your authored fail message, never a raw error.

Every check can carry an optional label, the short line learners read in that checklist ("Keeps only orders over €200"). Without one, the row says what the check looks at: "Your code", "Printed output", "Query result", or "Value of total". Labels are display copy only: adding or changing one never resets passes learners already earned.

YAML
checks:
  - type: sql_result_equals
    label: Keeps only shipped orders over €200
    rows: [[1041, 245.0], [1057, 612.5]]
    fail: Close. Keep the amount filter and add status = 'shipped'.
TypeLooks atPasses when
code_matchesthe editor codepattern (a regex) matches the code
stdout_containsprinted outputoutput contains text
stdout_matchesprinted outputpattern (a regex) matches the output
stdout_equalsprinted outputoutput equals text (trailing whitespace on each line and blank lines ignored; set trim: false for byte-exact)
sql_result_equalsthe query resultresult rows equal rows — optional columns: (names, case-insensitive) and orderMatters: true (by default row order is ignored, but duplicate rows still count)
global_equalsa Python variablethe global named name equals value (any JSON shape)

Which to use: SQL labs → sql_result_equals is the strongest check — it verifies what the query actually returned. Python labs → stdout_* for printed output, global_equals for computed variables. code_matches works in any language and is best as a shape hint alongside an outcome check.

Tip

The robust pattern is one code_matches for the concept ("uses a WHERE clause") plus one result/stdout check for the outcome. The shape check catches the right idea with a wrong result; the outcome check catches a right result reached the wrong way.

pattern and flags

pattern is a regular expression (JavaScript syntax, written without /…/ delimiters). flags is an optional string of regex flags:

  • i — ignore case (where matches WHERE)
  • m — ^ and $ match per line, not per whole text
  • s — . also matches newlines
  • u — full unicode matching

flags: i is the one you'll use most: learners shouldn't fail a check over capitalization. A useful token in patterns is \b (word boundary): '\bWHERE\b' matches the keyword WHERE but not everywhere.

Fail messages, hints, and the solution

  • fail is required on every check — publishing without one is impossible. Write it as help, not as an error: say what to do next ("There's no WHERE clause yet…"), not what went wrong internally. Fail messages and hints are plain text with one nicety: wrap a name or a snippet in backticks (`total`) and learners see it as code.
  • hints escalate: the first is free, each further failed attempt unlocks one more. Order them from nudge to near-answer.
  • solution powers "Apply the solution": after a failed attempt the learner can put it in the editor and move on. It applies known-good code and records a checkpoint (analytics separates clean passes from skips). Always write one — it's what keeps an unattended lab from hard-sticking someone, and it's what Test lab verifies against.

Important

Checkpoints unlock the next exercise, but on a course with an assessed certificate they don't count toward it. The learner can always re-attempt for a clean pass, and the course page tells them how many exercises still need one.

Passing finishes the lesson. Once every exercise of every lab in a lesson has a pass (checkpoints included) and every question of its quizzes is done, the lesson marks itself complete — learners never do the work and forget the button. Lessons with no labs or quizzes complete via Complete and continue.

Test your lab before publishing

Press Test lab on the lab's row under the editor (or ⌘K → "Test lab: …"). It walks the lab the way a learner would:

  • the code each exercise starts from — starterCode for the first, the previous exercise's solution after that — must fail every check. A check that already passes before the learner does anything is vacuous: a bug in the check, not the learner;
  • your solution must pass every check — otherwise the results say the check "fails on the authored solution", and "Apply the solution" would strand people.

The results show every authored fail message exactly as a failing learner sees it — proofread them there. A passing test marks the lab verified; any edit to the lab block flips it to Not tested since the last edit until you test again. Typos and structural mistakes never get that far: the editor lints the lab live as you type (unknown fields, unknown check types, patterns that don't compile), and errors block publishing.

Embedded labs

Some labs already exist outside the academy: a sandbox you host, a product playground, an exercise app your team built. An embedded lab frames that page inside a lesson, tells it who the learner is, and lets it report the result back, so a pass there counts exactly like a pass in a browser lab: the lesson can complete itself, the assessed certificate advances, lesson.completed fires.

The fence

lab
id: deploy-a-dag
title: 'Lab: Deploy your first DAG'
estimatedMinutes: 15
adapter: external
url: https://labs.acme.com/airflow/deploy
height: 640                 # optional, px (320–2000); the lab may resize itself
exercises:                  # optional — leave it out to treat the lab as one step
  - id: write
    title: Write the DAG
    steps:
      - md: In the lab below, create a DAG called `hello`.
  - id: deploy
    title: Deploy it

Rules the editor enforces as you type:

  • url must be https://, and its host must be trusted (next section) — an untrusted host is an error that blocks publishing.
  • language, setup, packages and starterCode don't apply (the page at url runs the code); exercises carry no checks, hints or solution (the lab judges the work). Instructions go on an exercise (brief, steps) and render above the frame; for a single step with instructions, list one exercise.
  • Without exercises, the lab is one step with the id lab. With them, the lab reports each step by its id, and from two exercises the bench shows the step strip. An empty list (exercises: []) is refused — leave the key out instead.

Test lab doesn't apply to embedded labs: there is nothing of ours to run. The author preview frames the lab with a preview token (below), so you can rehearse without recording anything; the token names you by your learner identity at the academy (opening the preview creates one if you don't have one yet).

Trusted hosts and the secret

Authors can't embed arbitrary pages. On Admin → Settings → Embedded labs, list the hosts you vouch for, one per line (labs.acme.com). The match is exact: trusting labs.acme.com doesn't trust acme.com or any other subdomain, so each host earns its own line.

The first save with a host mints the academy's embed secret. Keep it on your lab's server only: it verifies the launch token and signs the results. Rotate secret invalidates every token handed out so far; the lab's server must switch to the new value before it can report again.

Warning

Trust is checked live, not only at publish. Remove a host from the list and every published lesson that frames it stops at once: learners see "this lab isn't available right now" instead of the frame, and results from that host are refused (403) until the host is back on the list.

The launch token

When a signed-in learner reaches the lab, the frame's URL carries a mastria_token query parameter: a JWT (HS256, signed with the embed secret) saying who opened the lab and from where. The token is minted the moment the frame loads (as the bench scrolls into view), not when the page was rendered, so its 15-minute window starts when your page receives it; and the frame keeps that URL for the rest of the visit, so a result or a progress refresh never reloads your page. Verify the signature with any JWT library, then trust the claims:

ClaimMeaning
issThe academy's origin, e.g. https://academy.acme.com
audYour lab's origin — accept only tokens meant for it
subThe learner's stable id at this academy (the same id webhooks carry)
email, nameThe learner's sign-in address and display name (name may be null)
account{ id, name } of the company the learner belongs to, or null for individuals
tenant, course, lesson, labThe academy's slug and where the lab was opened — send the token back with the result
iat, expIssued at, and expiry 15 minutes later — the launch window
jtiA unique id per launch
previewtrue when staff open the lab from the editor's preview; show the lab, but expect no result to be recorded

Accept a token only while it's fresh, then run your own session. A signed-out visitor gets the frame without a token: show the lab if you like, and say that results count once they sign in on the academy. The bench says the same beside the frame.

Note

A learner id in a URL could be typed by anyone; a signed token can't be forged without the secret. That's the whole point of the token.

Reporting a result

When the learner passes (or fails) a step, your server posts:

POST https://mastria.dev/api/labs/external
Content-Type: application/json
X-Mastria-Signature: <hex HMAC-SHA256 of the raw body, keyed with the embed secret>

{
  "token": "<the launch token the lab received>",
  "exercise": "deploy",                 — omit for a single-step lab
  "passed": true,
  "checks": [                           — optional
    { "label": "DAG parses", "pass": true },
    { "label": "Task ran once", "pass": false, "message": "The task never ran — trigger the DAG." }
  ]
}

The signature is the same scheme as the outbound webhook's (hex(HMAC-SHA256(secret, rawBody))), so code that already verifies webhooks can sign results. The token names the learner, lesson and lab; Mastria verifies it with the academy's secret and accepts it for 7 days after the launch (its 15-minute exp is the launch window, not the result window). Preview tokens are refused, and so is a result for a learner the academy no longer allows on the course (a restriction added since the launch, or a sunset course they never started) — the same gate a browser lab's check passes through.

The response is { "recorded": true, "lessonCompleted": <bool> }, where lessonCompleted is true when this result completed the lesson. Errors come back with a plain-word error:

StatusMeaning
401The token or the signature doesn't verify
403A preview token, an untrusted host, or a learner no longer allowed on the course
404The course, lesson, lab or learner is gone or unpublished
409The lab id names a browser lab, not an embedded one
413The body is over 32 KB
422Invalid body, or an unknown or missing exercise

Why the body signature when the token already identifies the launch? A learner can read their own token out of the frame's URL. Without the server-side signature they could report their own pass; with it, only your server can.

A reported pass is a lab attempt like any other: the lesson completes when every lab and quiz is done, the assessed certificate counts it as a clean pass (anchored to the lab's URL and the step id, so moving the lab to another page starts the count over), and the lab shows up in course analytics. Failing checks are recorded there by their message (or label when there is no message); passing checks aren't itemized. The certificate's "labs author-verified" note speaks for the labs Mastria runs; embedded labs are outside that claim.

Talking to the page around you

The frame listens to postMessage from your lab's origin only:

  • { "type": "mastria:lab:resize", "height": 820 } — the frame takes that height (320–2000 px).
  • { "type": "mastria:lab:result" } — send it right after your server posted a result; the lesson re-reads the recorded state and the step ticks (or the "Lab complete" footer appears) without a reload. Nothing about the message itself is trusted — the recorded state is.

Learners also have a Refresh progress link beside the frame, for labs that never send the message.

Good to know

  • Pasting code into a | field indents it for you. Paste with the cursor right after setup: |, starterCode: | or solution: | (or on an empty line at the right depth), and every line lands at the depth the YAML needs, with the code's own indentation kept.
  • YAML errors name the line in your lesson, and the two classic traps come with the fix: a value that starts with a backtick, or text with a colon followed by a space, needs quotes around it.
  • One editor buffer per lab. starterCode seeds it, every exercise checks it, "Apply the solution" replaces it.
  • setup re-runs before every learner run, so your tables and variables always reset. Never write an exercise that depends on state a previous run created. Use idempotent DDL (CREATE OR REPLACE).
  • Learners can read setup. The lab shows it, folded away, under Defined for you above the editor, so the tables and variables your brief mentions don't appear from nowhere. Don't hide answers in it: it runs in the learner's browser either way. Checks and solutions are never shown.
  • Python packages (e.g. pandas) preload at startup; known packages a learner imports are also fetched automatically.
  • Keep estimatedMinutes honest — recalibrate from real completion times in analytics.
  • Numbers in sql_result_equals compare loosely (5 = 5.0 = "5"), so decimal formatting never fails an exact-value check.

During the design-partner phase, we author lab blocks with you — same-day, like all content.