Simulation¶

Use python/tatbot_sim/ for offline development, dataset-shape checks, and control experiments that do not connect to an arm or camera.

Ownership¶

Responsibility

Shared owner

Robot, tool, surface and sensor selection

Existing registries resolved into ResolvedConfig

Observation names, units and availability

tatbot_contracts and ObservationBuilder

Supported production reference motion

Cartesian compiler and C++ planner

Synthetic episode lifecycle

Episode, used by generation, evaluation and presentation

Physics, contacts and rendered sensors

ManiSkillWorld

Physical drawing runs on the ROS 2 stack in ros/ (tatbot ros draw); its own simulation is its hardware:=mock launch, described in ros/README.md. This package does not drive that stack. Dataset batching is an episode scheduler: it produces independent episodes that share world construction, observations and supported production references.

A future MuJoCo implementation should provide world construction, joint stepping, contact feedback and camera samples behind these boundaries. Keep its engine code and dependencies separate while reusing the robot/sensor registry, planner, observations and dataset writer. The adapter interface can still evolve when a second engine exercises it.

Quick start¶

cd python/tatbot_sim
uv sync --extra maniskill
uv run --extra maniskill python -m tatbot_sim.factory --list

Keep generated episodes and renders outside the repository. Include the source revision and simulator configuration in any artifact manifest.

Measured fixed-camera configurations require matching camera-bundle and robot-world calibration IDs. The simulator’s root is the follower arm base; the solver’s root is the full rig URDF. tatbot_sim.calibration composes that registration with the canonical arm mount before converting camera poses. The OpenCV optical axes (right, down, forward) are then converted to SAPIEN camera axes (forward, left, up). The fixed PoE camera body meshes are not part of the single-arm simulator model. The shared translation-only draw planner derives its mount offset from the URDF and refuses a rotated mount it cannot represent.

The full rig’s rig_center coincides with its root at the arm-pair midpoint; +Y is the robot’s own left side when looking forward through the fixed cameras. The single-arm simulator and measured palette poses remain in follower-base coordinates. A configured synthetic root-frame point is still an explicit development fixture, not an estimate of the physical rig midpoint. The simulator’s palette is the installed one, loaded from urdf/palette.urdf and config/palette.yaml (six caps), at the synthetic scene pose in config/palette_geometry.json (palette).

Posed-body work (sim compile, sim resolve, and their tests) resolves every anchor through the browser’s TypeScript InkLang resolver, so a sim host also needs Node.js 22 and the inkmap dependencies (npm ci in web/inkmap; scripts/check sim installs them when it can and otherwise reports a named skip). A user-local Node under ~/.local is enough.

The private engineering checkout reads its qualified workspace and arm profile. The public checkout instead falls back to the explicitly simulation-only files under config/examples/. They make imports and geometry-only development reproducible; they are not calibration, controller limits, or permission to run hardware. Offline generation continues with nominal geometry when a qualified calibration is unavailable, and records a development-only warning.

scripts/check sim and offline generation do not require a fresh tool calibration. When only nominal or synthetic geometry is available they run, stamp qualification: development, and retain the reason in geometry_warnings. Pass --require-qualified-geometry only when the artifact is explicitly meant to demonstrate calibrated contact geometry.

Explicit world construction¶

import tatbot_sim does not select a tool, register a Gym environment, build assets, or import the physics engine. Configuration and geometry helpers can be imported with only the standard library. Construct a world explicitly:

from tatbot_sim.resolved import resolve
from tatbot_sim.env import TatbotDrawEnv

config = resolve(tool_id="lutin-ballpoint-dot", seed=7)
world = TatbotDrawEnv(config=config, num_envs=1, obs_mode="rgbd",
                      control_mode="pd_joint_pos", sim_backend="cpu")
try:
    observation, info = world.reset(seed=7)
finally:
    world.close()

Resolution snapshots the existing tool, substrate, camera, calibration, timing, randomization and ink registries. Mutable values are copied at the boundary. Pass the same resolved configuration to the world, planner and IK solver; changing an environment variable later cannot change that world. Derived robot files are keyed by input content, and construction refuses changed source files after resolution. Factory and cinematic commands select their distribution’s tool explicitly and no longer restart Python.

Layout, camera mounting, lighting, action noise and RGB-D corruption use separate seeded streams. Generated tool metadata and policy evaluation results retain the resolved configuration and source digests. Engine reset seeds still identify individual episodes within a run.

Preview and cinematic cameras, lighting and surface appearance are instance options. Cinematic mounted views use separate cine_* cameras; they do not resize policy images. A historical lower-wrist shot requires the explicit historical camera profile. Ordinary construction uses TatbotDrawEnv directly; gym.make("TatbotDraw-v0") is no longer registered as an import side effect.

Shared episode runtime¶

Generation, policy evaluation, preview and cinematic rendering use Episode with ManiSkillWorld and ObservationBuilder. The runtime owns reset, articulation rebinding, material initialization, staged pose placement, reference refinement, control ticks and completion. Each command supplies its plan or policy actions and consumes the resulting observations. Every consumer shares the resolved world and production reference primitives; batching remains an episode-scheduling concern.

The observation contract lives in tatbot_contracts.observations. Its named seven joint positions and seven external-effort channels keep the existing LeRobot ordering. contact exposes the simulated contact estimate, with the resolved calibration when present; unavailable masks those channels to zero and records that they are unavailable. Policy evaluation explicitly uses the latter profile and clean RGB-D. Synthetic generation uses the contact profile and the recipe’s sensor corruption. These are recorded configuration choices. Simulator contact truth stays separate from policy features.

A tick creates one observation. Contact-feature reduction, saved depth and policy consumers reuse that sample, including its noise. Integer depth remains millimetres with zero invalid. The writer retains its existing log-depth codec. Pixel noise is keyed by camera, episode and control tick, so keeping every third video frame cannot change samples at the retained ticks. Presentation cameras remain outside the policy profile.

Runtime time starts at reset and advances by 1 / control_hz for each completed step. Dataset timestamps keep LeRobot’s zero-based first sample; metadata records the actual capture ticks and the sampling rule. Completion records whether the plan horizon or the backend ended the episode. A video truncated by its frame limit records an unfinished episode.

Episode metadata includes scene/reset seeds, actual camera mount poses and sampled lights, sensor response parameters and seeds, engine versions, and measurements of the final solved reference. Scene construction and episode placement use separate random streams. A scene retained across resets keeps its recorded construction seed; an explicit rebuilding reset redraws it from the requested seed. This preserves reproducibility without forcing reconstruction for every batched run.

Production reference primitives¶

The factory can consume the production Cartesian compiler and C++ joint planner. Select it with --production-draw-config PATH, pointing to an existing tatbot.draw-config/1 containing the tool and drawing speed:

tatbot sim generate paper-draw -- --out-dir /tmp/production-reference \
  --design design.json --production-draw-config draw.json \
  --no-tool-calibration-jitter --dr.latency.obs-delay-steps 0 0 \
  --num-episodes 1 --num-envs 1

This path currently supports nominal ballpoint geometry on an undisplaced paper plane, ordinary drawing intent and zero observation delay. The supplied draw configuration owns approach, speed, easing and carriage settings. The production compiler owns chunk limits, pen-up travel and reference motion. Native planner refusals remain refusals; the synthetic expert does not repair or perturb the accepted reference. Its IK is used only to propose the initial simulated pose and to measure the final reference. Simulation quality checks still apply.

Every native command executes at 400 Hz with three physics substeps (1200 Hz). The existing ManiSkill position controller follows the native position reference. Native velocity references remain in the retained JSON; this adapter does not apply their feed-forward term or claim vendor-controller equivalence. The configured 30 Hz camera captures on ticks 14, 27, 40, and so on: the first controller tick at or after its deadline, less than 2.5 ms late. No command or pen marker is resampled. LeRobot video/data retain their nominal 30 Hz grid; run_meta.json records each episode’s actual capture_control_ticks, and contact-feature derivatives use those actual sample times. The horizon flags retain their existing duration unit of 30 Hz frames. Episode steps_planned/steps_executed count controller ticks; frames_recorded and sidecar lengths count camera samples.

meta/references/ retains the surface, draw configuration, Cartesian CSVs and native joint-plan JSONs. Episode motion metadata identifies the planner, source digests, chunk boundaries, numeric conversion to float32, and absence of expert action noise. The plane and initial pose are explicitly simulated geometry.

This is reference-primitive reuse in the data factory. The factory path does not cover palette visits, and it does not claim that ink deposited at 400 Hz is numerically identical to the synthetic 30 Hz material model. Engine contact and material behavior require their own measurements. In the initial CPU replay, controller lag kept measured contact briefly during a commanded lift. The privileged timeline preserves that discrepancy and its existing audit reports it. A valid export and matching native references do not establish drawing-quality acceptance.

Camera profiles¶

deployment is the default sensor profile. It resolves physical arm assignment, RGB-D dimensions and cadence from the vision registry, with the checked-in example as the offline fallback. The follower environment renders the right arm’s one wrist view (wrist_upper in the current profile). It does not attach the left camera to the follower. Camera mounts come from the canonical URDF; intrinsics derived from nominal field of view are labeled nominal in metadata.

Use --sensor-profile legacy-two-view explicitly when reproducing historical datasets or checkpoints with wrist_upper and wrist_lower on the follower. This adds the historical lower camera geometry and does not describe the current installed robot. Generation, preview, and policy evaluation use the same selection. Presentation views never enter the dataset’s camera features.

tatbot sim generate paper-draw -- --out-dir /tmp/current-paper
tatbot sim generate paper-draw -- --out-dir /tmp/historical-paper \
  --sensor-profile legacy-two-view

Policy evaluation and no-arm wire probes compare the checkpoint’s exact image keys with the selected profile before querying actions. For a server-side checkpoint, pass --checkpoint-config /path/to/config.json with its local configuration. Simulation evaluation records that config’s digest. A missing view is an error; no image duplication or relabeling fills it.

Substrates¶

config/substrates.yaml is the one record of what the tools work on. The sim sizes its geometry and its texture from it, Inkmap’s Paper and Cylinder workspaces start from it, and the real workspace records its plane against the same numbers. Three substrates exist:

Substrate

Presentation

Size

Printed

paper_pad

flat pad, 10 mm thick

190.5 Ă— 279.4 mm (7.5 Ă— 11 in)

white, faint blue ÂĽ in (6.35 mm) square grid

paper_cylinder

rigid cylinder, ⌀85 mm

190.5 mm (7.5 in) long

the same grid all the way round

silicon_skin

flat sheet or wrapped

140 Ă— 185 mm

nothing

A tool datasheet names its default substrate and the others it admits: the ballpoint draws on either paper fixture, the laser and the 3RL only on the skin. TATBOT_SUBSTRATE=paper_cylinder selects an admitted alternative for a run; naming one the tool does not admit is refused rather than substituted. The paper cylinder’s canvas is the whole outer surface except the bottom quarter it rests on — three quarters of the circumference, 135° either side of the crest — along the full length of the cylinder; the bottom and the end caps are textured but never drawn on.

Material and surface profiles¶

For a sheet, material and shape are independent scenario axes. The paper-draw recipe uses flat paper by default, while both silicone recipes balance flat and cylinder members inside each vectorized batch. Sampled cylinders run along the long canvas direction and wrap the short direction at a 75-110 mm radius. A cylinder-shaped substrate is not sampled: the paper cylinder pins the profile to cylinder at its own 42.5 mm radius, and a balanced request on it is refused. Every episode records surface_profile and surface_radius_m.

Any sheet supports an explicit override:

tatbot sim generate paper-draw -- --out-dir /tmp/paper-cylinder \
  --dr.surface.profile cylinder
tatbot sim generate skin-tattoo -- --out-dir /tmp/skin-flat \
  --dr.surface.profile flat
TATBOT_SUBSTRATE=paper_cylinder tatbot sim generate paper-draw -- \
  --out-dir /tmp/paper-cylinder-fixture

balanced means both profiles occur in each batch with at least two environments. Curved visual geometry and the mathematical contact surface are generated from the same chart. Curved profiles remain kinematic-contact development data until their collision mesh is separately qualified.

A dataset to read without generating one¶

tatbot/sim-paper-draw-demo is a published shard of the paper-draw distribution — 8 episodes, 2836 frames, wrist RGB and depth, generated with seed 0 and no flags beyond the recipe. Use it to inspect the dataset shape, feature names and metadata that this repository produces before running the factory yourself:

uv run --project python/lerobot_robot_tatbot python -c "
from lerobot.datasets.lerobot_dataset import LeRobotDataset
ds = LeRobotDataset('tatbot/sim-paper-draw-demo')
print(ds.meta.info['total_episodes'], 'episodes;', ds.meta.info['total_frames'], 'frames')
print(sorted(ds.meta.features))"

With a locally qualified workspace and arm profile, regenerate an equivalent shard with:

tatbot sim generate paper-draw -- --out-dir <dir> --num-episodes 8 --num-envs 4 --seed 0

The public placeholder profile cannot reproduce that shard’s qualified contact geometry or calibration jitter. It can still generate development data using nominal tool dimensions, with the basis and warning recorded; do not relabel that output as calibrated evidence.

Each paper-draw invocation represents one fitted session: its seed chooses a single persistent tip offset within the recorded calibration uncertainty, and the run metadata records both the central calibration and the applied offset. Across shard seeds this varies plausible seating/calibration geometry while preserving exact agreement between the visible tip, TCP, collision, and marks.

Drawing a portable design¶

--design FILE draws one tatbot.inkmap-design/1 — a plane or cylinder chart design, as Inkmap’s Paper and Cylinder workspaces and tatbot design place export it — in place of the shared collection, for the artwork and erase tasks. The design’s strokes, rotation, mirror and scale are its own; the trajectory builder, the charge model and its dips, the reach mask, the cameras and the writer are the factory’s.

tatbot sim generate paper-draw -- --design design.json --out-dir <dir> \
  --num-episodes 1 --num-envs 1 --horizon 20000 --dr.ink.dips
TATBOT_SUBSTRATE=paper_cylinder tatbot sim generate paper-draw -- \
  --design cylinder-design.json --out-dir <dir> --horizon 20000

The design and the substrate have to be the same kind of surface: a plane design goes on a flat substrate as it is, a cylinder design on the paper cylinder of the same radius, turned a quarter so that Inkmap’s u (along the axis) becomes the canvas y the simulator runs along the axis and v (arc from the crest) becomes -x; both charts are right-handed about the outward normal, so the artwork keeps its handedness. A mismatch is refused, never bent to fit.

--design-placement authored (the default) keeps the anchor Inkmap saved, about the canvas centre, and refuses a placement the tool cannot be held normal to; sampled recentres the artwork and draws an offset inside the reach envelope per env, the way the collection is placed. A design is never shrunk to fit an episode: the run refuses up front, naming the --horizon its strokes need (a 70 × 75 mm filled design at a 0.3 mm hatch is roughly 3 m of line and 460 s at the ballpoint’s pacing, so the default 300 s artwork horizon will not hold it). The ink footprint is the design’s own stroke width. Every episode records the design’s digest and placement under artwork and program.portable_design, and the run metadata under design.

scripts/sim_preview.py --design FILE --task artwork renders the same plan from the wrist, third-person and top-down cameras without writing a dataset, and tatbot sim cinematic -- --portable-design FILE --task artwork shoots it path-traced, at social sizes, from the staged cameras aimed at the drawing.

The GPU physics backend also let a fixed-base articulation’s root creep (3.1 mm in 30 s on an idle arm, quadratic in time; none on the CPU backend). The environment now re-pins the robot root every control step, and python/tatbot_sim/tests/test_root_pin.py holds it there on any node with a CUDA device.

What simulation proves¶

Simulation can validate pure transforms, schema handling, deterministic replay, and software integration. It does not prove camera calibration, contact force, e-stop behavior on physical hardware, or safe human use.

Posed-body scenarios¶

Inkmap placement files describe design intent on a canonical rest-body surface. config/inkmap/tattoo-scenario.schema.json describes one resolved offline simulation realization: body and rig identity, pose, body/world and robot/world transforms, support fixture, declared tool, immutable SVG, and the derived face/barycentric stroke trace.

The editor also exports tatbot.inkmap-sim-bundle/1. sim compile recognizes this immutable request and emits the separate typed v3 scenario schema, binding the source bundle, InkProgram, and tool-profile digest. Multiple placements require --placement-id ID for drawing compilation. Pose/tool/seed/world overrides refuse rather than mutate the bundle. The v3 color/label renderer and compiled-preview import preserve the typed program and its surface binding. See design format for the source, validation, and asset rules.

InkLang is the only semantic placement resolver. It turns the description plus the fixed model/identity binding into a surface-bound resolution before scenario compilation. Simulation consumes that JSON through a thin Node.js 22 client; it does not parse InkLang or choose another site anchor in Python.

The split is intentional. A placement survives pose changes; a scenario is fully replayable. Inkmap region uv is semantic and normalized, so it must not be used as a metric drawing chart. The simulator maps the frozen metric program through a surface trace and then skins those face/barycentric points into the selected pose.

Named body poses are kinematic and static within an episode. This is a geometry and planning model, not a soft-tissue, breathing, or dynamic-person model. The checked-in SOMA GLB preserves the indexed mid-face address as an expanded render/picking view; its pose cache stores generated face vertices for every named session pose. Browser and Python loaders verify model, identity, topology, rest-surface, asset, and pose digests before using those bytes.

The optional mechanics ladder is separate. Rigid contact remains the admitted reference until synchronized, calibrated non-human phantom data qualify a compliant candidate. Synthetic mechanics fixtures exercise schemas, fitting, uncertainty, out-of-distribution refusal, solver energy, and gradients, but do not establish a physical contact model or authorize force/depth control.

Materialize one placement without launching SAPIEN:

tatbot sim compile config/inkmap/examples/forearm-placement-v6.json -- \
  --pose reclined-left-arm-supported --seed 42 --output /tmp/forearm-scenario.json

Compilation is CPU-only. It verifies body, rest-surface, rig, design, and placement identities before writing, and it fails explicitly on unsupported artwork constructs, non-manifold/open surface exits, or topology discontinuities. By default the patch normal is aligned to robot +Z and its +u axis uses the validated pi robot-world yaw; --target-world-m and --patch-yaw-rad expose that constrained body-to-robot placement. Dataset generation recomputes every trajectory’s FK and refuses the scenario if any target exceeds the 1 mm IK residual gate.

Materialize a deterministic coverage suite before launching the simulator:

tatbot sim sample --count 64 -- \
  --output-dir ~/tatbot-sim/scenarios/seed-42 --seed 42

A suite that is at least as large as its site list promises to cover every site, and that needs two slots per pose: with the default five poses, ask for at least 10 scenarios (or fewer than the six sites, which makes no coverage promise). Smaller requests fail up front with the minimum named instead of dying mid-run.

--sample N is the same module as compile (tatbot_sim.inkmap.cli sample); because it always builds the body-tattoo distribution, it declares the nominal lutin-3rl-bugpin tool independently of the workstation’s fitted or calibrated tool. An explicit contradictory --ee-tool is rejected. The nominal geometry is recorded as a development warning and never blocks this offline workflow on a calibration gate.

The sampler balances five tattoo-session poses (supine, prone, reclined with the legs on a chair rest, and reclined with either arm supported), and six initial atlas sites. By default it selects immutable artwork from the shared collection, using the training-family split. --artwork-split validation|test|all selects another explicit split. It reuses artwork across independently sampled scenes. The six sites are a simulation distribution subset, not the 59-site InkLang vocabulary. It constrains every compiled trace to the requested face-labeled region, pairs supported-arm poses only with a tattoo on that same arm, and selects an upward-exposed atlas face so the named pose stays meaningful relative to gravity. If one body/pose/site pairing fails clearance, the bounded retry advances to a different exposed site for that same balanced body/pose slot before repeating the pairing. A clearance rejection is retained and that exact physical pairing is deprioritized for later slots, while missing site coverage is tried first. It then searches the finite envelope in config/inkmap/placement-search.json. The CPU audit varies body X/Y, body yaw, and fixture offset; probes 32 exact trajectory targets; and rejects any candidate that misses the 1 mm IK, 10 mm non-tool, or 5 mm tool-shaft gates. The probe is a fast rejection gate only. The candidate that wins it is then solved in full, the way the generator solves it: the same planner over the whole 3600-step trajectory, the expert’s sequential joint solve, and FK of that joint reference against every target. A candidate whose full solve misses 1 mm anywhere is rejected as ik_full and the next-ranked candidate is tried, so an accepted scenario is one the generator will accept. (Before this gate, 2026-09-03, a placement passed 32 probes at 1e-7 m and then failed generate on 185 of 5400 targets that fell between them.) URDF collision meshes are sampled to a conservative 2 mm lower bound, while the body and fixture use named capsule/box proxies. Every selected scenario embeds the configuration digest, transforms, objective values, clearances, and complete candidate ledger. attempts.jsonl records suite rejects with explicit reasons; exhausting the bounded retry budget fails the suite. The later dataset run still recomputes FK for every target and refuses any residual over 1 mm. --no-reach-audit exists for geometry-only debugging; its output is not reach-qualified. To use acquired artwork, pass --design-source directory --generated-design-dir DIR, where DIR holds complete DrawingBot V3 acquisitions. Use --design-source spiral only for the fixed calibration spiral. The episode loop never makes a network request, and generated suites must be outside the checkout.

tatbot sim materialize (Inkgen) stores source imagery only: the exact PNG and traced SVG with their hashes and service metadata. The directory loader refuses it; acquire those sources through DBV3 first.

Resolve one semantic request through the canonical InkLang parser and bounded placement search:

tatbot sim resolve "a dbv3-orbit on the left forearm" -- \
  --output-dir ~/tatbot-sim/resolved/orbit-42 \
  --design-id dbv3-orbit --size-mm 30 30 --support armrest --seed 42

The typed tatbot.scenario-request/1 record captures the prompt, canonical placement intent, compatibility program, parser/config digest, design subject or exact ID, size, fixed model binding, site, side, pose, support, and seed. Free text cannot bypass InkLang validation or write scenario JSON. Normal requests select matching frozen collection artwork or complete immutable materializations; --design-id spiral-v1 is the sole built-in regression exception. Simulation invokes the same batch-capable TypeScript InkLang core as Inkmap and embeds its complete intent and resolution JSON in PlacementFile v6 provenance. Node.js 22 or newer is therefore an explicit simulation dependency; there is no second Python grammar or face resolver. A named simulation-exposed-grid-v1 policy may select only among canonically resolved region-UV candidates for the requested pose. Resolution produces request, placement, scenario, manifest, and attempt-ledger files. Unknown anatomy, missing laterality, incompatible pose/support, relative sites unsupported by this scenario consumer, and unsupported SVGs fail with named errors. The default provenance epoch keeps same-revision request/seed outputs byte-identical; pass --created-at when a wall-clock timestamp is part of the request provenance.

The compiled scenario enters the normal Tatbot expert, IK, floor-clamp, ink, render, and LeRobot writer through a separate distribution from the skin-tattoo silicone-pad scenes:

tatbot sim generate body-tattoo -- \
  --scenario /path/to/one/accepted.scenario.json \
  --out-dir ~/tatbot-sim/body-forearm --num-episodes 8 --num-envs 8 --seed 0

The scene includes the complete posed body as a kinematic visual, a textured drawable mesh patch, and conservative body-capsule clearance proxies. Chair, bed, armrest, and tabletop geometry is no longer instantiated, including furniture collisions and clearance obstacles. Pad scenes also omit tabletop clutter. Legacy support IDs remain readable as scenario pose metadata; legacy table and clutter randomization settings no longer create scene objects. Generated OBJ caches live under ~/.cache/tatbot/body-scenarios/; datasets remain outside the repository. Body-tattoo approach and inter-stroke hover are capped at 20 mm above the local surface; this preserves an approach while staying inside the audited 3RL orientation envelope. The drawable patch mesh reaches past the pigment field’s raster, so its texture is sampled with edge clamping; a repeating sampler used to tile the drawn design across the whole limb in every wrist camera.

The current body patch is a curved, kinematically projected contact surface; the coarse body capsules are avoidance proxies, not a qualified skin contact mesh. Body datasets are therefore stamped kinematic-contact-v1 and target the resolved TCP at zero working offset. The 3RL currently uses nominal datasheet geometry: generation, preview, compile, audit, and evaluation continue with a machine-readable development warning instead of waiting on a touch-off. Use --require-qualified-geometry only when the purpose of a run is to prove calibration eligibility. An axis-inferred body is not itself a warning when its axisymmetric contact tool has a qualified fixed-point calibration.

Exact-design simulation evaluation¶

Ask generation to score each deposition episode while the exact intended path, surface, and pigment field are still in memory:

tatbot sim generate paper-draw -- \
  --out-dir ~/tatbot-sim/eval/expert-seed-4200 \
  --num-episodes 8 --num-envs 8 --seed 4200 --judge
tatbot sim eval dataset ~/tatbot-sim/eval/expert-seed-4200 -- \
  --output-dir ~/tatbot-sim/eval/reports/expert-4200 \
  --training-seed-range 0 4095

Every batch, for every distribution, is gated on the joint reference the expert actually solved: FK of that reference must land within 1 mm of every target, and on every pen-down step it must sit within 0.25 mm of its intended height above the surface. The second gate exists because the damped IK trades position against the requested tool lean and its converged reference sat 0.5-2 mm above the sheet on later strokes: inside the residual gate, outside the 0.5 mm contact band, so the pen hovered and never marked. Generation now closes the reference on contact, moving each offending target along the surface normal by its measured error and re-solving (up to three rounds). A batch that still misses is re-solved once with a longer budget, and an episode that still misses is dropped, as is an episode whose sheet nothing touched. Each episode records its final reference errors in run_meta. Dropped episodes are never written; they are listed in run_meta.dropped_episodes with their reason (ik_reference, dip_reference or idle) and the run keeps generating until the requested count is met. Pass --keep-idle to retain blank episodes for a study of them. Before this gate the sequential solve could drift centimetres off a stroke unnoticed: a paper stroke sat 38 mm above the sheet with the run reporting no fault (2026-09-03).

A dip is gated on the cap rather than on its commanded point. The point is inside the cap by construction, so the 1 mm residual cannot tell a dip from a reference that never entered one — and the charge is credited by step index, so nothing downstream notices either. Generation now checks, at each credited step, that the tip lies within that cap’s own radius of the cap axis and that the tool is within 15 degrees of the entry axis; an episode that misses is dropped as dip_reference. Measure a placement before generating dip episodes with tatbot sim reach, which reports both per cap. Before this gate, every cap of the palette installed then, at its synthetic placement, sat 5-25 mm and 17-26 degrees outside, and full charges were credited for dips that never reached a cap (2026-09-09).

Typed Inkmap v3 scenarios additionally materialize target-labels.npz, target-coverage.png, target-reference.png, and a convention/hash manifest beside the posed scene geometry. These are canonical chart-space intent at 4 px/mm, not runtime state: soft coverage is area sampled, colors are unassociated sRGB, and integer layer/placement IDs are nearest sampled with zero as background. Surface-chart rotation and texture mirror follow the same split as the browser, avoiding a second baked rotation. The body patch starts at the bundle’s requested skin tone; only contact-driven InkField deposition changes runtime pigment. This avoids leaking a completed target tattoo into policy observations.

The reproducible renderer-parity command is:

uv run --project python/tatbot_sim --extra maniskill \
  python -m tatbot_sim.inkmap.parity_evidence --output /absolute/new/evidence-dir

It records accepted/rejected denominators, exact surface-address and posed position discrepancies, connected-domain/exclusion/semantic-region leakage, complete-wrap refusal, aligned mask IoU, boundary distance, original/browser/ simulator views, and difference images. Its 12 px/mm comparison raster is an adequate-resolution qualification view, not a change to the 4 px/mm minimum target emitted with ordinary scenarios.

Render a strict CPU perception reference from a typed v3 scenario:

tatbot sim perception /tmp/scenario-v3.json -- \
  --output-dir /tmp/inkmap-perception --views 3 --seed 42

Each frame stores RGB, clean and declared-corrupted metric depth with separate validity, camera-frame geometric normals, visible-body mask, global SOMA face index and perspective-correct barycentrics, and visible tattoo soft coverage, placement ID, and layer ID. Integer IDs are nearest sampled and zero means no tattoo; non-body pixels use face -1 and zero barycentrics/normals. Occluders retain valid scene depth while clearing body and tattoo semantics. Calibration, independent per-axis seeds, sampled appearance/camera/depth/occlusion settings, source and asset provenance, identity/design/placement/scenario hashes, and a leakage-safe split are in each manifest. These NPZ sidecars are privileged labels and are not policy observation features.

Re-run the fail-closed audit with python -m tatbot_sim.inkmap.perception_audit --path DIR. The CPU path is a semantic oracle, not an assumed production renderer: its report includes frames/second, peak memory, bytes/frame, and linear 540-frame projections. Do not scale until an assigned renderer meets the recorded throughput/storage budget and human review accepts the images. The reference identity is admitted; three bounded MHR/SOMA candidates remain refused for dataset use until their human visual-review cells change from pending. Their topology alone is not transfer authorization.

Plan the complete pilot before assigning compute:

tatbot sim pilot-plan -- --output-dir /tmp/inkmap-pilot --seed 42
tatbot sim pilot-audit /tmp/inkmap-pilot/pilot-plan.json

This starts no renderer and materializes no scenes. It writes a hash-bound, audited ledger for 60 reference-identity scenario templates, the selected three-identity/540-view expansion, 12 reach-gated drawing episodes, all named start/failure variants, and the separate stress/refusal suite. Candidate identity, compute-host, reach/clearance, GPU, and human-review gates remain machine-readable and pending rather than being inferred from the plan. Stencil and failure injection are implemented contracts; only generated evidence for them remains gated by compute, reach, rendering, and review.

For a generated drawing episode, --save-privileged-labels writes an NPZ timeline under meta/privileged rather than adding policy observation fields. It requires --texture-refresh-steps 1; otherwise generation refuses before launch, because a new deposited-ink label paired with stale RGB is not the same simulation step. Each retained episode records post-step tool pose, synthetic contact distance/incidence and pen state, intended target/surface frame, compiled primitive and TattooProgram layer, elapsed progress, deposited coverage, remaining target, and a synchronization bit. The dataset auditor checks hashes, one timeline per episode, exact frame count, target/contact consistency, monotonic progress, bounded fractions, and synchronization. Contact and deposition remain outputs of the declared synthetic model—not measured force, penetration, tissue response, or human-contact evidence.

Pass --episode-variant with one of blank-start, stencil-start, missed-stroke, interrupted, partial-coverage, dry-tool, or occluded when generating a compiled body scenario. Stencil pixels remain a separate appearance and privileged label below deposited pigment: the per-step stencil_visible_fraction is the mean of stencil coverage times one minus deposited pigment, so it falls as the drawing covers its guide and the auditor refuses a timeline where it grows; the constant guide area is stencil_area_fraction in run_meta. An occluded episode hides exactly --occlusion-fraction of every camera image (the pilot ledger carries the value from its appearance) and records the achieved area. Failure variants keep the unmodified intended trajectory as their answer key, change only the named execution or observation dimension by a declared constant (6 mm tangent slide, 15 mm lift over the middle 40 % of steps, stop at 55 % of intended steps), retain otherwise-idle episodes, and write an explicit expected outcome rather than a successful-demonstration label. The plan’s reach masks and tool ceiling are validated on the unperturbed targets, so the lift is checked against the same ceiling and the joint-reference gate is re-run on the perturbed targets; a variant that leaves the IK envelope is refused with a message naming the variant, never blamed on the compiled tattoo.

The judge rasterizes the answer key with the same InkField kernel that lays down simulated pigment. Its headline is tolerance-band F1: precision measures whether drawn pigment belongs to the design, while recall measures how much of the design was completed. IoU, directional Chamfer distances, coverage ratio, blank/engaged fractions, contact duration, interaction frames, and floor-clamp rates explain that number. Corruption tests pin the expected direction for offset, scale error, truncation, jitter, and missing strokes.

Every judged episode stores exact intended, drawn, and overlay PNGs with hashes. tatbot sim eval dataset writes a JSON report, scores.csv, and a human-readable Markdown report, including a deterministic 95% bootstrap interval when at least three episodes exist. A training-seed overlap marks the split contaminated; dirty dataset source, dirty evaluator source, mixed tools/distributions, or fewer than three episodes makes the report non-comparable. Producer and checkpoint identity come from the dataset contract and cannot be relabeled by the report command.

A clean held-out result is still reported as screen-only. Simulation ranking has not yet established correlation with physical rollouts, and no score from this command is motion or human-contact authorization.

Run a checkpoint closed-loop through the same async LeRobot server used by a rollout:

scripts/eval/serve.sh --policy /models/candidate --env-root ~/il-serve
tatbot sim eval policy -- \
  --server 127.0.0.1:8080 --policy /models/candidate \
  --wire-scenario act_rgbd14_masked --distribution paper-draw \
  --repetitions 3 --seed 20260903 \
  --output-dir ~/tatbot-sim/eval/candidate-20260903

A blank-sheet control runs the same worker with no server and no checkpoint:

tatbot sim eval policy -- \
  --client-mode hold-control --distribution paper-draw \
  --repetitions 3 --seed 20260910 --output-dir ~/tatbot-sim/eval/hold-20260910

It sends the worker’s own joint state back every step, scores F1 0, and runs from the simulator’s interpreter, so a sim node without the serving environment can produce it. Its report carries the checkpoint id hold-control and a digest of that contract name, never a model digest.

The LeRobot client and ManiSkill worker remain separate processes and virtual environments. Their local socket carries bounded JSON plus typed arrays, while the client deliberately uses the deployed gRPC inference protocol. Feature keys and shapes come from the actual Tatbot follower declaration. Chunk overlap, stale-timestep rejection, 30 Hz target filtering, joint slew, worker protocol, checkpoint digest, zero-effort basis, surface profile, tool geometry basis/warnings, intended/drawn/overlay hashes, and optional videos are retained in the episode bundle. --resume reuses only completed deterministic episodes.

The fixed spiral is reserved for regression controls. Normal policy batteries use the seed-generated design stream (or explicitly acquired DBV3 artwork), so a successful client cannot overfit a small checked-in design catalog.

Contribution checklist¶

  1. Add a deterministic fixture for the new behavior.

  2. Assert units and coordinate frames at the boundary.

  3. Run scripts/check --light. Run scripts/check sim where a locally qualified simulation profile is present; otherwise preserve its explicit profile-missing skip and exercise the config-independent tests you changed.

  4. Label simulated results as simulated in the run manifest.

Shared artwork and calibration¶

Normal body, paper and skin drawing use dbv3-acquired-v1, the three native DBV3 acquisitions in Inkmap’s manifest. The train, validation and test splits hold out their source families. Each acquisition has one physical size and pen configuration; resizing or changing pen width requires regeneration. Acquired paths retain order and direction through the simulator schedule. The spiral remains an explicit simulator calibration control, never a finished artwork fallback. --generated-design-dir now reads native acquisition directories with result.json, artwork.json and a matching frozen recipe. Older traced-image materializations require new DBV3 acquisition.

tatbot sim sample --count 12 -- --output-dir ~/tatbot-sim/artwork-train \
  --artwork-split train
tatbot sim sample --count 4 -- --output-dir ~/tatbot-sim/artwork-test \
  --artwork-split test --poses reclined-left-arm-supported --sites forearm
tatbot sim generate paper-draw -- --out-dir ~/tatbot-sim/paper-art \
  --num-episodes 4 --artwork-split train
tatbot sim qualify-artwork -- --output-dir ~/tatbot-sim/artwork-review
tatbot sim qualify-artwork -- --output-dir ~/tatbot-sim/artwork-review-hatch --fill-style hatch

qualify-artwork plans the collection through the same scheduled material strokes the planner schedules for a drawing; --fill-style selects the paint planner (fill styles) and the report records it, so the two planners can be compared artwork by artwork on stroke count, ideal footprint IoU and simulated deposition.

Recipes¶

tatbot sim recipes expands a frozen artwork library into reproducible scenario recipes without compiling, rendering or executing anything:

tatbot sim recipes -- --output-dir ~/tatbot-sim/recipes-v1 --count 1000 --seed 7
tatbot sim recipes-status -- --output-dir ~/tatbot-sim/recipes-v1
# native DBV3 acquisitions instead of the bundled collection:
tatbot sim recipes -- --output-dir ~/tatbot-sim/recipes-gen --count 1000 \
  --artwork-dir ~/tatbot-artwork/dbv3-acquisitions

A recipe is a pure function of the frozen plan and its index. Every axis — artwork, identity, pose, site, scale, rotation, mirror, target pose, and the named camera/appearance/sensor/compile streams a later stage will use — is drawn from its own seed derived from (plan seed, recipe key, axis name). So sharding the work, resuming it, or reordering it produces byte-identical recipes, and the plan’s digest is the run’s identity. Rerunning the command resumes: recipes already on disk are verified against the plan and kept, anything that no longer matches is rewritten from the plan rather than trusted.

The ledger counts requested, admitted, rejected, compiled, rendered and executed separately, and the last three stay zero here. A recipe is not a render, and a render is not a completed drawing. Rejected recipes keep their reason in rejected.jsonl.

Splits are assigned from the artwork’s family and the identity, before any augmentation — never from a variant’s seed, crop or SVG digest — so a rotated or rescaled copy of a training artwork cannot become held-out artwork. The command audits family and identity leakage across the held-out partitions and exits non-zero if it finds any.

Artwork with no reviewed family — a freshly generated library, for instance — is one group, so the whole of it moves across splits together. That is the conservative reading of unknown provenance rather than a meaningful family split, and the ledger names the artwork in unknown_family_artworks instead of leaving it to be inferred from a single-split tally. Families are reviewed metadata; nothing here invents one from a subject line.

The collection binds exact acquired JSON hashes, source/license metadata, recipe identity, generation width, frozen physical size, and artwork families. The three bundled families have one training, one validation and one test example. Placement rotation and scene variation retain the family’s split; duplicates cannot cross it. A different physical size requires regeneration through native DBV3. Episode metadata records source hashes and family/split. Evaluation reports block artwork from training or unknown splits and report calibration controls separately from artwork comparisons. Historical datasets are not relabeled.

Normal sampling and semantic resolution compile typed v3 bundles. The placement optimizer authors new immutable bundle requests for each candidate, binding body position, yaw, and bounded support offset. Its candidate/rejection ledger lives in the suite manifest; it does not mutate an already compiled v3 world transform. All existing IK and robot/tool-shaft clearance gates remain.

Paper/skin artwork defaults to a 9,000-step (five-minute) cap at 30 Hz — config.ARTWORK_HORIZON_STEPS, raised from 3,600 on 2026-09-10. Complete artworks must fit; randomizing an artwork never drops a stroke to meet the cap. The cap is not free: a candidate that used to be refused cheaply at two minutes now runs the full IK reach audit over three times the trajectory, so a suite that leaves the audit on takes correspondingly longer. --no-reach-audit skips the CPU yaw selection; final generation still enforces exact FK.

tatbot sim sample -- --max-seconds N bounds a suite’s wall clock (default 1800, 0 for none). The budget is enforced in two places because one is not enough: between candidates, and as a share of the remaining budget around each placement search. Wrapping only the last stage of that search left a measured run twenty minutes past a ten-minute budget without the check ever being reached. A search that outruns its share is refused with reason time_budget and recorded like any other rejection; the suite then finishes with the honest partial report it already produces for an incomplete run. The default footprint is a nominal simulated 0.3 mm, not measured tool calibration. V3 episodes use their bound uniform footprint width; mixed-width episodes refuse until per-stroke width execution is available. Continuous contact is sampled between control frames so a narrow line does not become disconnected dots; pen lifts and resets break that interpolation.

What the planner will and will not draw¶

Two geometric gates decide whether artwork becomes a drawing, and both moved on 2026-09-10 at the fleet owner’s direction. Neither is a safety limit: they gate fidelity, not motion, contact force, retract or ink accounting.

Paint coverage (fill_geometry.PAINT_COVERAGE_FLOOR, 0.90, was 0.98) is how much of a painted region the planned stroke path must actually cover. The old threshold refused 57 of 95 candidates in a measured suite over generated flash. Those refusals are not one population: their coverage runs 0.0, 0.31, 0.77 … 0.97, and a floor of 0.70 admits exactly the same set as a floor of 0.50, because a handful cover essentially nothing. 0.90 leaves a tenth of the ink at most missing, and a drawing that covers nothing is still refused. It costs little against a looser floor: the same library gives 9 usable artworks at 0.98, 11 at 0.90 and 12 at 0.80. A consumer that wants the old strictness passes coverage_floor=0.98.

Paint thinner than the tool used to refuse the whole drawing. A tool that cannot draw a line thinner than its own tip does not refuse to draw the line: it traces the middle and the line comes out at tool width, which is what a person does with a fine liner. Those regions are now traced down the middle, and ink is allowed outside the artwork only in the halo the tool needs for them. How much heavier the result may be is bounded by NARROW_OVERDRAW_LIMIT (4x the painted area): a 0.3 mm tool on a 0.2 mm line lays 1.76x and is drawn; on a 0.05 mm hair it would lay far more and is refused, as is any tool larger than the region itself.

Measured on a 24-piece generated library at 30-45 mm: 11 of 24 compile to strokes, against 9 before these changes. Of the twelve that do not, four carry paint finer than the tool can meaningfully trace and five cover essentially nothing at any size — properties of what the model drew. Generate roughly twice the artwork a suite needs and expect to discard about half.

Qualification compares source SVG, shared typed preview, canonical target, ideal finite-width coverage, complete expert trajectory duration, and runtime InkField deposition. The reference uses 48 px/mm with 4x target supersampling; paint comparisons require IoU >=0.98 and runtime deposition requires >=0.90. The latter is a discrete nominal pigment model, separate from source-paint parity. Each size, rejection, duration, stroke count, and pen lift is recorded. Output includes a visual contact sheet and source/compiled/target/deposition images. This CPU planar qualification does not establish GPU scene, robot IK, measured deposition, or physical execution acceptance. The body sampler and bounded generated episodes provide separate evidence for their own stages.

Artwork generation resolves pigment and pad textures at 16 px/mm (the flat recipe can override --artwork-pixels-per-m). Typed body patches use the same 16 px/mm pigment minimum; reference targets retain their independent minimum. The higher resolution costs GPU memory, so batch size must fit the selected node. Nominal width is 0.3 mm, independent of the original 2–4 mm legacy ink DR. Raster imports use the shared polygon tracer, retaining small connected marks and holes; spline fitting is excluded because it can distort paint without reporting an error.

Regenerate the five-pose artwork gallery with tatbot sim showcase-artwork -- --output-dir /path/outside/checkout. Review and install its six JSON files in web/inkmap/public/showcase/; its manifest preserves offline-only qualification. The line and geometry files under config/inkmap/examples/ remain contract fixtures, not normal design sources or gallery content.

Dependency and check boundaries¶

The default tatbot_sim install provides CPU preparation, geometry and dataset utilities. Its numerical dependencies include PyTorch; it does not install ManiSkill or SAPIEN. Use uv sync --project python/tatbot_sim --extra maniskill on a simulation host. Generation, preview and cinematic launchers select this extra explicitly. Policy evaluation with a separately provisioned worker interpreter requires the same extra in that interpreter.

tatbot check sim runs the partitions below. tatbot check sim-fast selects the same partitions and excludes marked slow integration cases. Neither runs in the push hook. Each missing capability reports its own SKIP; it cannot skip contracts or count as full coverage.

Check

Scope

Additional requirements

sim-contracts

Resolved configuration, observation/transport contracts, CPU imports

Base Python environment

sim-geometry

Numerical geometry, planning and design compilation

Node with TypeScript support, C++ toolchain, cached robot assets

sim-engine

ManiSkill adapter and material/control contracts

Engine extra, cached robot assets and compiler toolchain

sim-render

Rendered sensors, shared episodes, dataset round trips

Engine extra, render device and compiler toolchain

Individual groups can run with tatbot check sim-contracts, for example. Slow rendering cases need the C++ planner, a render device and the checkout’s retained tool calibration; missing toolchains and the public nominal profile report explicit skips. sim-fast excludes them. This is developer regression coverage and adds no hardware-execution gate. The public profile builds its compiler fixtures only for groups that use them. Engine assets retain ManiSkill’s cache layout; reading the cache path or constructing texture files does not import the physics engine.