dyno lab LOCAL PREVIEWApp guide
Development preview · not released

Compare model responses with a saved research protocol.

Write the question and grading criteria first. Run matched conditions, inspect the answers, and keep a record of each attempt. This local preview adds controlled comparisons to Dyno Lab’s existing research notebooks.

01PrepareDefine the prompt, conditions, rubric and request budget.
02RunSelect a running model. Save each request and response.
03ReviewCompare matched answers and record your labels.
04ReproduceExport evidence or prepare a linked local reproduction.

A real application check

We ran two locally authored cases with neutral and user-pressure instructions on Qwen3.8-27B-MLX-4bit. All four requests returned final answers. The screenshot shows those saved responses in the native app. They remain ungraded.

Dyno Lab controlled studies showing saved policy and geography responses side by side
Local software acceptance run, September 16, 2026. Two cases, two conditions, seed 0, thinking off, temperature 0.3, maximum 192 output tokens per request. This checks the workflow. It is not a benchmark replication, a statistical study or a safety score.

Review answers before revealing their conditions

Saved responses can now be reviewed in a shuffled queue. The shared prompt, rubric and answer remain visible; condition labels, model identity and other reviewers’ judgments are hidden. Record prior exposure honestly. Reveal is explicit and permanent for that review session.

Native review queue with one saved judgment and condition labels hidden
Software acceptance using a saved real response. The reviewer is named “UI acceptance (automated)” and reports prior exposure. This demonstrates the interface, not independent human validation.

Read the app walkthrough and SDK/API/MCP examples. Context masking is a review aid, not access control. Answer wording may disclose context. Bounded monitor evaluations are now available in this local preview.

Compare a monitor with saved judgments

Choose answer-only, thinking-only or combined evidence. Freeze the reference judgments and threshold before running. The report keeps development and test results separate and shows missing evidence, invalid scores and disagreements.

Native monitor report with two scored development responses and unavailable recall
Real 27B workflow check with two automated reference judgments, both passes. There are no positive examples, so recall is unavailable. This tests execution and reporting, not monitor quality.

Monitor setup, SDK and API instructions.

Test probes with separate data and baseline controls

Keep related examples in one partition. Fit on training data, choose the layer and regularization on validation data, and inspect the test result alongside shuffled-label, majority and text-length controls. Intervention comparisons can also include no-op, restored baseline and magnitude-matched random directions.

Native grouped probe report showing held-out scores and baseline controls
A 12-example negation fixture on the downloaded 27B model, with four test examples. This is a software acceptance check, not a new interpretability or safety finding.

Run grouped probes and intervention controls.

Try the workflow in the local preview

  1. Start a model in Models, or a compatible generation endpoint in Pools.
  2. Open Lab → Studies → Controlled comparisons.
  3. Enter your question, hypothesis, rubric and two instructions. Set the sample count, thinking option and token budget.
  4. Save the protocol, select the running endpoint and run the conditions.
  5. Review each answer against the rubric. Failed, cancelled and truncated attempts remain separate.
  6. Export evidence. Importing evidence does not run it; prepare a reproduction when you want a new local run.

The quick editor creates one development case. Multi-case protocols and separate development/test groups are available through the local API and Python SDK. MCP separates protocol preparation from execution.

Six additional workflows, ready for local review

The app, API, SDK and MCP changes are implemented locally. Community changes use a separate isolated database for validation. Nothing on this page is a new scientific finding.

FeatureWhat it doesGuide
Behavioral regression reportsPaired changes, missing judgments and group bootstrap intervals.Step by step
Artifact compatibilityRevision, tokenizer, architecture, layer, pooling and shape checks; unknown metadata blocks eligibility.Step by step
Monitor threshold selectionSelect on development scores, preserve the choice and disclose known test exposure.Step by step
Community reproductionsParent-version links, explicit comparison criteria, evidence references and private drafts.Step by step
Simulated agent tasksBounded in-memory actions, saved observations and separate correctness/disclosure checks.Step by step
Checkpoint comparisonsMatching-protocol reports with declared checkpoint, adapter and training-data provenance.Step by step
Dyno Lab regression-reports with saved acceptance results
A paired report from actual exact-token software acceptance responses. This is not a safety evaluation.
Dyno Lab compatibility-checks with saved acceptance results
The newly trained 27B grouped probe compared with its own recorded contract. This checks metadata eligibility, not cross-model transfer.
Dyno Lab checkpoint-comparisons with saved acceptance results
Two actual runs of the same checkpoint and protocol. The zero differences are an acceptance check, not an improvement claim.
Dyno Lab simulated-agent-tasks with saved acceptance results
Actual 27B model actions and final answer. The test runner is deliberately simulated as unavailable.
Dyno Lab monitor-evaluations with saved acceptance results
Actual 27B monitor responses on two development and two test examples. The examples are an exact-token software fixture, not a monitor benchmark.

Limits: compatibility checks do not apply artifacts or attest weight integrity. Simulated tasks cannot execute real code. Checkpoint comparisons do not train models or verify supplied lineage. The community migration is local and awaits a reviewed deployment.

Keep conclusions separate from measurements

Saved outputs let other people inspect what happened. They do not establish that a model is safe, that a probe identifies a causal mechanism, or that an intervention generalizes. The next development stages focus on controls, reproducibility and clearer limits on what each result supports.

This page and its screenshots are local review material. No release or production website update has been published.