Compare model responses with a saved research protocol.
Write the question and grading criteria first. Run matched conditions, inspect the answers, and keep a record of each attempt. This local preview adds controlled comparisons to Dyno Lab’s existing research notebooks.
A real application check
We ran two locally authored cases with neutral and user-pressure instructions on Qwen3.8-27B-MLX-4bit. All four requests returned final answers. The screenshot shows those saved responses in the native app. They remain ungraded.

Review answers before revealing their conditions
Saved responses can now be reviewed in a shuffled queue. The shared prompt, rubric and answer remain visible; condition labels, model identity and other reviewers’ judgments are hidden. Record prior exposure honestly. Reveal is explicit and permanent for that review session.

Read the app walkthrough and SDK/API/MCP examples. Context masking is a review aid, not access control. Answer wording may disclose context. Bounded monitor evaluations are now available in this local preview.
Compare a monitor with saved judgments
Choose answer-only, thinking-only or combined evidence. Freeze the reference judgments and threshold before running. The report keeps development and test results separate and shows missing evidence, invalid scores and disagreements.

Monitor setup, SDK and API instructions.
Test probes with separate data and baseline controls
Keep related examples in one partition. Fit on training data, choose the layer and regularization on validation data, and inspect the test result alongside shuffled-label, majority and text-length controls. Intervention comparisons can also include no-op, restored baseline and magnitude-matched random directions.

Run grouped probes and intervention controls.
Try the workflow in the local preview
- Start a model in Models, or a compatible generation endpoint in Pools.
- Open Lab → Studies → Controlled comparisons.
- Enter your question, hypothesis, rubric and two instructions. Set the sample count, thinking option and token budget.
- Save the protocol, select the running endpoint and run the conditions.
- Review each answer against the rubric. Failed, cancelled and truncated attempts remain separate.
- Export evidence. Importing evidence does not run it; prepare a reproduction when you want a new local run.
The quick editor creates one development case. Multi-case protocols and separate development/test groups are available through the local API and Python SDK. MCP separates protocol preparation from execution.
Six additional workflows, ready for local review
The app, API, SDK and MCP changes are implemented locally. Community changes use a separate isolated database for validation. Nothing on this page is a new scientific finding.
| Feature | What it does | Guide |
|---|---|---|
| Behavioral regression reports | Paired changes, missing judgments and group bootstrap intervals. | Step by step |
| Artifact compatibility | Revision, tokenizer, architecture, layer, pooling and shape checks; unknown metadata blocks eligibility. | Step by step |
| Monitor threshold selection | Select on development scores, preserve the choice and disclose known test exposure. | Step by step |
| Community reproductions | Parent-version links, explicit comparison criteria, evidence references and private drafts. | Step by step |
| Simulated agent tasks | Bounded in-memory actions, saved observations and separate correctness/disclosure checks. | Step by step |
| Checkpoint comparisons | Matching-protocol reports with declared checkpoint, adapter and training-data provenance. | Step by step |





Limits: compatibility checks do not apply artifacts or attest weight integrity. Simulated tasks cannot execute real code. Checkpoint comparisons do not train models or verify supplied lineage. The community migration is local and awaits a reviewed deployment.
Keep conclusions separate from measurements
Saved outputs let other people inspect what happened. They do not establish that a model is safe, that a probe identifies a causal mechanism, or that an intervention generalizes. The next development stages focus on controls, reproducibility and clearer limits on what each result supports.
This page and its screenshots are local review material. No release or production website update has been published.