--- summary: Hardware-independent Tatbot simulation workflow tags: [simulation, testing] updated: 2026-09-27 audience: [dev, contributor] --- # Simulation Use `python/tatbot_sim/` for offline development, dataset-shape checks, and control experiments that do not connect to an arm or camera. ## Ownership | Responsibility | Shared owner | | --- | --- | | Robot, tool, surface and sensor selection | Existing registries resolved into `ResolvedConfig` | | Observation names, units and availability | `tatbot_contracts` and `ObservationBuilder` | | Supported production reference motion | Cartesian compiler and C++ planner | | Synthetic episode lifecycle | `Episode`, used by generation, evaluation and presentation | | Physics, contacts and rendered sensors | `ManiSkillWorld` | Physical drawing runs on the ROS 2 stack in `ros/` (`tatbot ros draw`); its own simulation is its `hardware:=mock` launch, described in `ros/README.md`. This package does not drive that stack. Dataset batching is an episode scheduler: it produces independent episodes that share world construction, observations and supported production references. A future MuJoCo implementation should provide world construction, joint stepping, contact feedback and camera samples behind these boundaries. Keep its engine code and dependencies separate while reusing the robot/sensor registry, planner, observations and dataset writer. The adapter interface can still evolve when a second engine exercises it. ## Quick start ```bash cd python/tatbot_sim uv sync --extra maniskill uv run --extra maniskill python -m tatbot_sim.factory --list ``` Keep generated episodes and renders outside the repository. Include the source revision and simulator configuration in any artifact manifest. Measured fixed-camera configurations require matching camera-bundle and robot-world calibration IDs. The simulator's root is the follower arm base; the solver's root is the full rig URDF. `tatbot_sim.calibration` composes that registration with the canonical arm mount before converting camera poses. The OpenCV optical axes (right, down, forward) are then converted to SAPIEN camera axes (forward, left, up). The fixed PoE camera body meshes are not part of the single-arm simulator model. The shared translation-only draw planner derives its mount offset from the URDF and refuses a rotated mount it cannot represent. The full rig's `rig_center` coincides with its root at the arm-pair midpoint; +Y is the robot's own left side when looking forward through the fixed cameras. The single-arm simulator and measured palette poses remain in follower-base coordinates. A configured synthetic root-frame point is still an explicit development fixture, not an estimate of the physical rig midpoint. The simulator's palette is the installed one, loaded from `urdf/palette.urdf` and `config/palette.yaml` (six caps), at the synthetic scene pose in `config/palette_geometry.json` ([palette](palette.md)). Posed-body work (`sim compile`, `sim resolve`, and their tests) resolves every anchor through the browser's TypeScript InkLang resolver, so a sim host also needs Node.js 22 and the inkmap dependencies (`npm ci` in `web/inkmap`; `scripts/check sim` installs them when it can and otherwise reports a named skip). A user-local Node under `~/.local` is enough. The private engineering checkout reads its qualified workspace and arm profile. The public checkout instead falls back to the explicitly simulation-only files under `config/examples/`. They make imports and geometry-only development reproducible; they are not calibration, controller limits, or permission to run hardware. Offline generation continues with nominal geometry when a qualified calibration is unavailable, and records a development-only warning. `scripts/check sim` and offline generation do not require a fresh tool calibration. When only nominal or synthetic geometry is available they run, stamp `qualification: development`, and retain the reason in `geometry_warnings`. Pass `--require-qualified-geometry` only when the artifact is explicitly meant to demonstrate calibrated contact geometry. ## Explicit world construction `import tatbot_sim` does not select a tool, register a Gym environment, build assets, or import the physics engine. Configuration and geometry helpers can be imported with only the standard library. Construct a world explicitly: ```python from tatbot_sim.resolved import resolve from tatbot_sim.env import TatbotDrawEnv config = resolve(tool_id="lutin-ballpoint-dot", seed=7) world = TatbotDrawEnv(config=config, num_envs=1, obs_mode="rgbd", control_mode="pd_joint_pos", sim_backend="cpu") try: observation, info = world.reset(seed=7) finally: world.close() ``` Resolution snapshots the existing tool, substrate, camera, calibration, timing, randomization and ink registries. Mutable values are copied at the boundary. Pass the same resolved configuration to the world, planner and IK solver; changing an environment variable later cannot change that world. Derived robot files are keyed by input content, and construction refuses changed source files after resolution. Factory and cinematic commands select their distribution's tool explicitly and no longer restart Python. Layout, camera mounting, lighting, action noise and RGB-D corruption use separate seeded streams. Generated tool metadata and policy evaluation results retain the resolved configuration and source digests. Engine reset seeds still identify individual episodes within a run. Preview and cinematic cameras, lighting and surface appearance are instance options. Cinematic mounted views use separate `cine_*` cameras; they do not resize policy images. A historical lower-wrist shot requires the explicit historical camera profile. Ordinary construction uses `TatbotDrawEnv` directly; `gym.make("TatbotDraw-v0")` is no longer registered as an import side effect. ## Shared episode runtime Generation, policy evaluation, preview and cinematic rendering use `Episode` with `ManiSkillWorld` and `ObservationBuilder`. The runtime owns reset, articulation rebinding, material initialization, staged pose placement, reference refinement, control ticks and completion. Each command supplies its plan or policy actions and consumes the resulting observations. Every consumer shares the resolved world and production reference primitives; batching remains an episode-scheduling concern. The observation contract lives in `tatbot_contracts.observations`. Its named seven joint positions and seven external-effort channels keep the existing LeRobot ordering. `contact` exposes the simulated contact estimate, with the resolved calibration when present; `unavailable` masks those channels to zero and records that they are unavailable. Policy evaluation explicitly uses the latter profile and clean RGB-D. Synthetic generation uses the contact profile and the recipe's sensor corruption. These are recorded configuration choices. Simulator contact truth stays separate from policy features. A tick creates one observation. Contact-feature reduction, saved depth and policy consumers reuse that sample, including its noise. Integer depth remains millimetres with zero invalid. The writer retains its existing log-depth codec. Pixel noise is keyed by camera, episode and control tick, so keeping every third video frame cannot change samples at the retained ticks. Presentation cameras remain outside the policy profile. Runtime time starts at reset and advances by `1 / control_hz` for each completed step. Dataset timestamps keep LeRobot's zero-based first sample; metadata records the actual capture ticks and the sampling rule. Completion records whether the plan horizon or the backend ended the episode. A video truncated by its frame limit records an unfinished episode. Episode metadata includes scene/reset seeds, actual camera mount poses and sampled lights, sensor response parameters and seeds, engine versions, and measurements of the final solved reference. Scene construction and episode placement use separate random streams. A scene retained across resets keeps its recorded construction seed; an explicit rebuilding reset redraws it from the requested seed. This preserves reproducibility without forcing reconstruction for every batched run. ### Production reference primitives The factory can consume the production Cartesian compiler and C++ joint planner. Select it with `--production-draw-config PATH`, pointing to an existing `tatbot.draw-config/1` containing the tool and drawing speed: ```bash tatbot sim generate paper-draw -- --out-dir /tmp/production-reference \ --design design.json --production-draw-config draw.json \ --no-tool-calibration-jitter --dr.latency.obs-delay-steps 0 0 \ --num-episodes 1 --num-envs 1 ``` This path currently supports nominal ballpoint geometry on an undisplaced paper plane, ordinary drawing intent and zero observation delay. The supplied draw configuration owns approach, speed, easing and carriage settings. The production compiler owns chunk limits, pen-up travel and reference motion. Native planner refusals remain refusals; the synthetic expert does not repair or perturb the accepted reference. Its IK is used only to propose the initial simulated pose and to measure the final reference. Simulation quality checks still apply. Every native command executes at 400 Hz with three physics substeps (1200 Hz). The existing ManiSkill position controller follows the native position reference. Native velocity references remain in the retained JSON; this adapter does not apply their feed-forward term or claim vendor-controller equivalence. The configured 30 Hz camera captures on ticks 14, 27, 40, and so on: the first controller tick at or after its deadline, less than 2.5 ms late. No command or pen marker is resampled. LeRobot video/data retain their nominal 30 Hz grid; `run_meta.json` records each episode's actual `capture_control_ticks`, and contact-feature derivatives use those actual sample times. The horizon flags retain their existing duration unit of 30 Hz frames. Episode `steps_planned`/`steps_executed` count controller ticks; `frames_recorded` and sidecar lengths count camera samples. `meta/references/` retains the surface, draw configuration, Cartesian CSVs and native joint-plan JSONs. Episode motion metadata identifies the planner, source digests, chunk boundaries, numeric conversion to float32, and absence of expert action noise. The plane and initial pose are explicitly simulated geometry. This is reference-primitive reuse in the data factory. The factory path does not cover palette visits, and it does not claim that ink deposited at 400 Hz is numerically identical to the synthetic 30 Hz material model. Engine contact and material behavior require their own measurements. In the initial CPU replay, controller lag kept measured contact briefly during a commanded lift. The privileged timeline preserves that discrepancy and its existing audit reports it. A valid export and matching native references do not establish drawing-quality acceptance. ## Camera profiles `deployment` is the default sensor profile. It resolves physical arm assignment, RGB-D dimensions and cadence from the vision registry, with the checked-in example as the offline fallback. The follower environment renders the right arm's one wrist view (`wrist_upper` in the current profile). It does not attach the left camera to the follower. Camera mounts come from the canonical URDF; intrinsics derived from nominal field of view are labeled nominal in metadata. Use `--sensor-profile legacy-two-view` explicitly when reproducing historical datasets or checkpoints with `wrist_upper` and `wrist_lower` on the follower. This adds the historical lower camera geometry and does not describe the current installed robot. Generation, preview, and policy evaluation use the same selection. Presentation views never enter the dataset's camera features. ```bash tatbot sim generate paper-draw -- --out-dir /tmp/current-paper tatbot sim generate paper-draw -- --out-dir /tmp/historical-paper \ --sensor-profile legacy-two-view ``` Policy evaluation and no-arm wire probes compare the checkpoint's exact image keys with the selected profile before querying actions. For a server-side checkpoint, pass `--checkpoint-config /path/to/config.json` with its local configuration. Simulation evaluation records that config's digest. A missing view is an error; no image duplication or relabeling fills it. ## Substrates `config/substrates.yaml` is the one record of what the tools work on. The sim sizes its geometry and its texture from it, Inkmap's Paper and Cylinder workspaces start from it, and the real workspace records its plane against the same numbers. Three substrates exist: | Substrate | Presentation | Size | Printed | | --- | --- | --- | --- | | `paper_pad` | flat pad, 10 mm thick | 190.5 × 279.4 mm (7.5 × 11 in) | white, faint blue ¼ in (6.35 mm) square grid | | `paper_cylinder` | rigid cylinder, ⌀85 mm | 190.5 mm (7.5 in) long | the same grid all the way round | | `silicon_skin` | flat sheet or wrapped | 140 × 185 mm | nothing | A tool datasheet names its default substrate and the others it admits: the ballpoint draws on either paper fixture, the laser and the 3RL only on the skin. `TATBOT_SUBSTRATE=paper_cylinder` selects an admitted alternative for a run; naming one the tool does not admit is refused rather than substituted. The paper cylinder's canvas is the whole outer surface except the bottom quarter it rests on — three quarters of the circumference, 135° either side of the crest — along the full length of the cylinder; the bottom and the end caps are textured but never drawn on. ## Material and surface profiles For a sheet, material and shape are independent scenario axes. The `paper-draw` recipe uses flat paper by default, while both silicone recipes balance `flat` and `cylinder` members inside each vectorized batch. Sampled cylinders run along the long canvas direction and wrap the short direction at a 75-110 mm radius. A cylinder-shaped substrate is not sampled: the paper cylinder pins the profile to `cylinder` at its own 42.5 mm radius, and a `balanced` request on it is refused. Every episode records `surface_profile` and `surface_radius_m`. Any sheet supports an explicit override: ```bash tatbot sim generate paper-draw -- --out-dir /tmp/paper-cylinder \ --dr.surface.profile cylinder tatbot sim generate skin-tattoo -- --out-dir /tmp/skin-flat \ --dr.surface.profile flat TATBOT_SUBSTRATE=paper_cylinder tatbot sim generate paper-draw -- \ --out-dir /tmp/paper-cylinder-fixture ``` `balanced` means both profiles occur in each batch with at least two environments. Curved visual geometry and the mathematical contact surface are generated from the same chart. Curved profiles remain kinematic-contact development data until their collision mesh is separately qualified. ## A dataset to read without generating one [`tatbot/sim-paper-draw-demo`](https://huggingface.co/datasets/tatbot/sim-paper-draw-demo) is a published shard of the `paper-draw` distribution — 8 episodes, 2836 frames, wrist RGB and depth, generated with seed 0 and no flags beyond the recipe. Use it to inspect the dataset shape, feature names and metadata that this repository produces before running the factory yourself: ```bash uv run --project python/lerobot_robot_tatbot python -c " from lerobot.datasets.lerobot_dataset import LeRobotDataset ds = LeRobotDataset('tatbot/sim-paper-draw-demo') print(ds.meta.info['total_episodes'], 'episodes;', ds.meta.info['total_frames'], 'frames') print(sorted(ds.meta.features))" ``` With a locally qualified workspace and arm profile, regenerate an equivalent shard with: ```bash tatbot sim generate paper-draw -- --out-dir --num-episodes 8 --num-envs 4 --seed 0 ``` The public placeholder profile cannot reproduce that shard's qualified contact geometry or calibration jitter. It can still generate development data using nominal tool dimensions, with the basis and warning recorded; do not relabel that output as calibrated evidence. Each `paper-draw` invocation represents one fitted session: its seed chooses a single persistent tip offset within the recorded calibration uncertainty, and the run metadata records both the central calibration and the applied offset. Across shard seeds this varies plausible seating/calibration geometry while preserving exact agreement between the visible tip, TCP, collision, and marks. ## Drawing a portable design `--design FILE` draws one `tatbot.inkmap-design/1` — a plane or cylinder chart design, as Inkmap's Paper and Cylinder workspaces and [`tatbot design place`](design.md) export it — in place of the shared collection, for the artwork and erase tasks. The design's strokes, rotation, mirror and scale are its own; the trajectory builder, the charge model and its dips, the reach mask, the cameras and the writer are the factory's. ```bash tatbot sim generate paper-draw -- --design design.json --out-dir \ --num-episodes 1 --num-envs 1 --horizon 20000 --dr.ink.dips TATBOT_SUBSTRATE=paper_cylinder tatbot sim generate paper-draw -- \ --design cylinder-design.json --out-dir --horizon 20000 ``` The design and the substrate have to be the same kind of surface: a plane design goes on a flat substrate as it is, a cylinder design on the paper cylinder of the same radius, turned a quarter so that Inkmap's `u` (along the axis) becomes the canvas `y` the simulator runs along the axis and `v` (arc from the crest) becomes `-x`; both charts are right-handed about the outward normal, so the artwork keeps its handedness. A mismatch is refused, never bent to fit. `--design-placement authored` (the default) keeps the anchor Inkmap saved, about the canvas centre, and refuses a placement the tool cannot be held normal to; `sampled` recentres the artwork and draws an offset inside the reach envelope per env, the way the collection is placed. A design is never shrunk to fit an episode: the run refuses up front, naming the `--horizon` its strokes need (a 70 × 75 mm filled design at a 0.3 mm hatch is roughly 3 m of line and 460 s at the ballpoint's pacing, so the default 300 s artwork horizon will not hold it). The ink footprint is the design's own stroke width. Every episode records the design's digest and placement under `artwork` and `program.portable_design`, and the run metadata under `design`. `scripts/sim_preview.py --design FILE --task artwork` renders the same plan from the wrist, third-person and top-down cameras without writing a dataset, and `tatbot sim cinematic -- --portable-design FILE --task artwork` shoots it path-traced, at social sizes, from the staged cameras aimed at the drawing. The GPU physics backend also let a fixed-base articulation's root creep (3.1 mm in 30 s on an idle arm, quadratic in time; none on the CPU backend). The environment now re-pins the robot root every control step, and `python/tatbot_sim/tests/test_root_pin.py` holds it there on any node with a CUDA device. ## What simulation proves Simulation can validate pure transforms, schema handling, deterministic replay, and software integration. It does not prove camera calibration, contact force, e-stop behavior on physical hardware, or safe human use. ## Posed-body scenarios Inkmap placement files describe design intent on a canonical rest-body surface. `config/inkmap/tattoo-scenario.schema.json` describes one resolved offline simulation realization: body and rig identity, pose, body/world and robot/world transforms, support fixture, declared tool, immutable SVG, and the derived face/barycentric stroke trace. The editor also exports `tatbot.inkmap-sim-bundle/1`. `sim compile` recognizes this immutable request and emits the separate typed v3 scenario schema, binding the source bundle, InkProgram, and tool-profile digest. Multiple placements require `--placement-id ID` for drawing compilation. Pose/tool/seed/world overrides refuse rather than mutate the bundle. The v3 color/label renderer and compiled-preview import preserve the typed program and its surface binding. See [design format](design-format.md) for the source, validation, and asset rules. [InkLang](inklang.md) is the only semantic placement resolver. It turns the description plus the fixed model/identity binding into a surface-bound resolution before scenario compilation. Simulation consumes that JSON through a thin Node.js 22 client; it does not parse InkLang or choose another site anchor in Python. The split is intentional. A placement survives pose changes; a scenario is fully replayable. Inkmap region `uv` is semantic and normalized, so it must not be used as a metric drawing chart. The simulator maps the frozen metric program through a surface trace and then skins those face/barycentric points into the selected pose. Named body poses are kinematic and static within an episode. This is a geometry and planning model, not a soft-tissue, breathing, or dynamic-person model. The checked-in SOMA GLB preserves the indexed mid-face address as an expanded render/picking view; its pose cache stores generated face vertices for every named session pose. Browser and Python loaders verify model, identity, topology, rest-surface, asset, and pose digests before using those bytes. The optional mechanics ladder is separate. Rigid contact remains the admitted reference until synchronized, calibrated non-human phantom data qualify a compliant candidate. Synthetic mechanics fixtures exercise schemas, fitting, uncertainty, out-of-distribution refusal, solver energy, and gradients, but do not establish a physical contact model or authorize force/depth control. Materialize one placement without launching SAPIEN: ```bash tatbot sim compile config/inkmap/examples/forearm-placement-v6.json -- \ --pose reclined-left-arm-supported --seed 42 --output /tmp/forearm-scenario.json ``` Compilation is CPU-only. It verifies body, rest-surface, rig, design, and placement identities before writing, and it fails explicitly on unsupported artwork constructs, non-manifold/open surface exits, or topology discontinuities. By default the patch normal is aligned to robot +Z and its +u axis uses the validated `pi` robot-world yaw; `--target-world-m` and `--patch-yaw-rad` expose that constrained body-to-robot placement. Dataset generation recomputes every trajectory's FK and refuses the scenario if any target exceeds the 1 mm IK residual gate. Materialize a deterministic coverage suite before launching the simulator: ```bash tatbot sim sample --count 64 -- \ --output-dir ~/tatbot-sim/scenarios/seed-42 --seed 42 ``` A suite that is at least as large as its site list promises to cover every site, and that needs two slots per pose: with the default five poses, ask for at least 10 scenarios (or fewer than the six sites, which makes no coverage promise). Smaller requests fail up front with the minimum named instead of dying mid-run. `--sample N` is the same module as `compile` (`tatbot_sim.inkmap.cli sample`); because it always builds the `body-tattoo` distribution, it declares the nominal `lutin-3rl-bugpin` tool independently of the workstation's fitted or calibrated tool. An explicit contradictory `--ee-tool` is rejected. The nominal geometry is recorded as a development warning and never blocks this offline workflow on a calibration gate. The sampler balances five tattoo-session poses (supine, prone, reclined with the legs on a chair rest, and reclined with either arm supported), and six initial atlas sites. By default it selects immutable artwork from the shared collection, using the training-family split. `--artwork-split validation|test|all` selects another explicit split. It reuses artwork across independently sampled scenes. The six sites are a simulation distribution subset, not the 59-site InkLang vocabulary. It constrains every compiled trace to the requested face-labeled region, pairs supported-arm poses only with a tattoo on that same arm, and selects an upward-exposed atlas face so the named pose stays meaningful relative to gravity. If one body/pose/site pairing fails clearance, the bounded retry advances to a different exposed site for that same balanced body/pose slot before repeating the pairing. A clearance rejection is retained and that exact physical pairing is deprioritized for later slots, while missing site coverage is tried first. It then searches the finite envelope in `config/inkmap/placement-search.json`. The CPU audit varies body X/Y, body yaw, and fixture offset; probes 32 exact trajectory targets; and rejects any candidate that misses the 1 mm IK, 10 mm non-tool, or 5 mm tool-shaft gates. The probe is a fast rejection gate only. The candidate that wins it is then solved in full, the way the generator solves it: the same planner over the whole 3600-step trajectory, the expert's sequential joint solve, and FK of that joint reference against every target. A candidate whose full solve misses 1 mm anywhere is rejected as `ik_full` and the next-ranked candidate is tried, so an accepted scenario is one the generator will accept. (Before this gate, 2026-09-03, a placement passed 32 probes at 1e-7 m and then failed `generate` on 185 of 5400 targets that fell between them.) URDF collision meshes are sampled to a conservative 2 mm lower bound, while the body and fixture use named capsule/box proxies. Every selected scenario embeds the configuration digest, transforms, objective values, clearances, and complete candidate ledger. `attempts.jsonl` records suite rejects with explicit reasons; exhausting the bounded retry budget fails the suite. The later dataset run still recomputes FK for every target and refuses any residual over 1 mm. `--no-reach-audit` exists for geometry-only debugging; its output is not reach-qualified. To use acquired artwork, pass `--design-source directory --generated-design-dir DIR`, where `DIR` holds complete [DrawingBot V3](drawingbot.md) acquisitions. Use `--design-source spiral` only for the fixed calibration spiral. The episode loop never makes a network request, and generated suites must be outside the checkout. `tatbot sim materialize` (Inkgen) stores source imagery only: the exact PNG and traced SVG with their hashes and service metadata. The directory loader refuses it; acquire those sources through DBV3 first. Resolve one semantic request through the canonical InkLang parser and bounded placement search: ```bash tatbot sim resolve "a dbv3-orbit on the left forearm" -- \ --output-dir ~/tatbot-sim/resolved/orbit-42 \ --design-id dbv3-orbit --size-mm 30 30 --support armrest --seed 42 ``` The typed `tatbot.scenario-request/1` record captures the prompt, canonical placement intent, compatibility program, parser/config digest, design subject or exact ID, size, fixed model binding, site, side, pose, support, and seed. Free text cannot bypass InkLang validation or write scenario JSON. Normal requests select matching frozen collection artwork or complete immutable materializations; `--design-id spiral-v1` is the sole built-in regression exception. Simulation invokes the same batch-capable TypeScript InkLang core as Inkmap and embeds its complete intent and resolution JSON in PlacementFile v6 provenance. Node.js 22 or newer is therefore an explicit simulation dependency; there is no second Python grammar or face resolver. A named `simulation-exposed-grid-v1` policy may select only among canonically resolved region-UV candidates for the requested pose. Resolution produces request, placement, scenario, manifest, and attempt-ledger files. Unknown anatomy, missing laterality, incompatible pose/support, relative sites unsupported by this scenario consumer, and unsupported SVGs fail with named errors. The default provenance epoch keeps same-revision request/seed outputs byte-identical; pass `--created-at` when a wall-clock timestamp is part of the request provenance. The compiled scenario enters the normal Tatbot expert, IK, floor-clamp, ink, render, and LeRobot writer through a separate distribution from the `skin-tattoo` silicone-pad scenes: ```bash tatbot sim generate body-tattoo -- \ --scenario /path/to/one/accepted.scenario.json \ --out-dir ~/tatbot-sim/body-forearm --num-episodes 8 --num-envs 8 --seed 0 ``` The scene includes the complete posed body as a kinematic visual, a textured drawable mesh patch, and conservative body-capsule clearance proxies. Chair, bed, armrest, and tabletop geometry is no longer instantiated, including furniture collisions and clearance obstacles. Pad scenes also omit tabletop clutter. Legacy support IDs remain readable as scenario pose metadata; legacy table and clutter randomization settings no longer create scene objects. Generated OBJ caches live under `~/.cache/tatbot/body-scenarios/`; datasets remain outside the repository. Body-tattoo approach and inter-stroke hover are capped at 20 mm above the local surface; this preserves an approach while staying inside the audited 3RL orientation envelope. The drawable patch mesh reaches past the pigment field's raster, so its texture is sampled with edge clamping; a repeating sampler used to tile the drawn design across the whole limb in every wrist camera. The current body patch is a curved, kinematically projected contact surface; the coarse body capsules are avoidance proxies, not a qualified skin contact mesh. Body datasets are therefore stamped `kinematic-contact-v1` and target the resolved TCP at zero working offset. The 3RL currently uses nominal datasheet geometry: generation, preview, compile, audit, and evaluation continue with a machine-readable development warning instead of waiting on a touch-off. Use `--require-qualified-geometry` only when the purpose of a run is to prove calibration eligibility. An axis-inferred body is not itself a warning when its axisymmetric contact tool has a qualified fixed-point calibration. ## Exact-design simulation evaluation Ask generation to score each deposition episode while the exact intended path, surface, and pigment field are still in memory: ```bash tatbot sim generate paper-draw -- \ --out-dir ~/tatbot-sim/eval/expert-seed-4200 \ --num-episodes 8 --num-envs 8 --seed 4200 --judge tatbot sim eval dataset ~/tatbot-sim/eval/expert-seed-4200 -- \ --output-dir ~/tatbot-sim/eval/reports/expert-4200 \ --training-seed-range 0 4095 ``` Every batch, for every distribution, is gated on the joint reference the expert actually solved: FK of that reference must land within 1 mm of every target, and on every pen-down step it must sit within 0.25 mm of its intended height above the surface. The second gate exists because the damped IK trades position against the requested tool lean and its converged reference sat 0.5-2 mm above the sheet on later strokes: inside the residual gate, outside the 0.5 mm contact band, so the pen hovered and never marked. Generation now closes the reference on contact, moving each offending target along the surface normal by its measured error and re-solving (up to three rounds). A batch that still misses is re-solved once with a longer budget, and an episode that still misses is dropped, as is an episode whose sheet nothing touched. Each episode records its final `reference` errors in `run_meta`. Dropped episodes are never written; they are listed in `run_meta.dropped_episodes` with their reason (`ik_reference`, `dip_reference` or `idle`) and the run keeps generating until the requested count is met. Pass `--keep-idle` to retain blank episodes for a study of them. Before this gate the sequential solve could drift centimetres off a stroke unnoticed: a paper stroke sat 38 mm above the sheet with the run reporting no fault (2026-09-03). A dip is gated on the cap rather than on its commanded point. The point is inside the cap by construction, so the 1 mm residual cannot tell a dip from a reference that never entered one — and the charge is credited by step index, so nothing downstream notices either. Generation now checks, at each credited step, that the tip lies within that cap's own radius of the cap axis and that the tool is within 15 degrees of the entry axis; an episode that misses is dropped as `dip_reference`. Measure a placement before generating dip episodes with `tatbot sim reach`, which reports both per cap. Before this gate, every cap of the palette installed then, at its synthetic placement, sat 5-25 mm and 17-26 degrees outside, and full charges were credited for dips that never reached a cap (2026-09-09). Typed Inkmap v3 scenarios additionally materialize `target-labels.npz`, `target-coverage.png`, `target-reference.png`, and a convention/hash manifest beside the posed scene geometry. These are canonical chart-space intent at 4 px/mm, not runtime state: soft coverage is area sampled, colors are unassociated sRGB, and integer layer/placement IDs are nearest sampled with zero as background. Surface-chart rotation and texture mirror follow the same split as the browser, avoiding a second baked rotation. The body patch starts at the bundle's requested skin tone; only contact-driven `InkField` deposition changes runtime pigment. This avoids leaking a completed target tattoo into policy observations. The reproducible renderer-parity command is: ```bash uv run --project python/tatbot_sim --extra maniskill \ python -m tatbot_sim.inkmap.parity_evidence --output /absolute/new/evidence-dir ``` It records accepted/rejected denominators, exact surface-address and posed position discrepancies, connected-domain/exclusion/semantic-region leakage, complete-wrap refusal, aligned mask IoU, boundary distance, original/browser/ simulator views, and difference images. Its 12 px/mm comparison raster is an adequate-resolution qualification view, not a change to the 4 px/mm minimum target emitted with ordinary scenarios. Render a strict CPU perception reference from a typed v3 scenario: ```bash tatbot sim perception /tmp/scenario-v3.json -- \ --output-dir /tmp/inkmap-perception --views 3 --seed 42 ``` Each frame stores RGB, clean and declared-corrupted metric depth with separate validity, camera-frame geometric normals, visible-body mask, global SOMA face index and perspective-correct barycentrics, and visible tattoo soft coverage, placement ID, and layer ID. Integer IDs are nearest sampled and zero means no tattoo; non-body pixels use face `-1` and zero barycentrics/normals. Occluders retain valid scene depth while clearing body and tattoo semantics. Calibration, independent per-axis seeds, sampled appearance/camera/depth/occlusion settings, source and asset provenance, identity/design/placement/scenario hashes, and a leakage-safe split are in each manifest. These NPZ sidecars are privileged labels and are not policy observation features. Re-run the fail-closed audit with `python -m tatbot_sim.inkmap.perception_audit --path DIR`. The CPU path is a semantic oracle, not an assumed production renderer: its report includes frames/second, peak memory, bytes/frame, and linear 540-frame projections. Do not scale until an assigned renderer meets the recorded throughput/storage budget and human review accepts the images. The reference identity is admitted; three bounded MHR/SOMA candidates remain refused for dataset use until their human visual-review cells change from pending. Their topology alone is not transfer authorization. Plan the complete pilot before assigning compute: ```bash tatbot sim pilot-plan -- --output-dir /tmp/inkmap-pilot --seed 42 tatbot sim pilot-audit /tmp/inkmap-pilot/pilot-plan.json ``` This starts no renderer and materializes no scenes. It writes a hash-bound, audited ledger for 60 reference-identity scenario templates, the selected three-identity/540-view expansion, 12 reach-gated drawing episodes, all named start/failure variants, and the separate stress/refusal suite. Candidate identity, compute-host, reach/clearance, GPU, and human-review gates remain machine-readable and pending rather than being inferred from the plan. Stencil and failure injection are implemented contracts; only generated evidence for them remains gated by compute, reach, rendering, and review. For a generated drawing episode, `--save-privileged-labels` writes an NPZ timeline under `meta/privileged` rather than adding policy observation fields. It requires `--texture-refresh-steps 1`; otherwise generation refuses before launch, because a new deposited-ink label paired with stale RGB is not the same simulation step. Each retained episode records post-step tool pose, synthetic contact distance/incidence and pen state, intended target/surface frame, compiled primitive and TattooProgram layer, elapsed progress, deposited coverage, remaining target, and a synchronization bit. The dataset auditor checks hashes, one timeline per episode, exact frame count, target/contact consistency, monotonic progress, bounded fractions, and synchronization. Contact and deposition remain outputs of the declared synthetic model—not measured force, penetration, tissue response, or human-contact evidence. Pass `--episode-variant` with one of `blank-start`, `stencil-start`, `missed-stroke`, `interrupted`, `partial-coverage`, `dry-tool`, or `occluded` when generating a compiled body scenario. Stencil pixels remain a separate appearance and privileged label below deposited pigment: the per-step `stencil_visible_fraction` is the mean of stencil coverage times one minus deposited pigment, so it falls as the drawing covers its guide and the auditor refuses a timeline where it grows; the constant guide area is `stencil_area_fraction` in `run_meta`. An `occluded` episode hides exactly `--occlusion-fraction` of every camera image (the pilot ledger carries the value from its appearance) and records the achieved area. Failure variants keep the unmodified intended trajectory as their answer key, change only the named execution or observation dimension by a declared constant (6 mm tangent slide, 15 mm lift over the middle 40 % of steps, stop at 55 % of intended steps), retain otherwise-idle episodes, and write an explicit expected outcome rather than a successful-demonstration label. The plan's reach masks and tool ceiling are validated on the unperturbed targets, so the lift is checked against the same ceiling and the joint-reference gate is re-run on the perturbed targets; a variant that leaves the IK envelope is refused with a message naming the variant, never blamed on the compiled tattoo. The judge rasterizes the answer key with the same `InkField` kernel that lays down simulated pigment. Its headline is tolerance-band F1: precision measures whether drawn pigment belongs to the design, while recall measures how much of the design was completed. IoU, directional Chamfer distances, coverage ratio, blank/engaged fractions, contact duration, interaction frames, and floor-clamp rates explain that number. Corruption tests pin the expected direction for offset, scale error, truncation, jitter, and missing strokes. Every judged episode stores exact intended, drawn, and overlay PNGs with hashes. `tatbot sim eval dataset` writes a JSON report, `scores.csv`, and a human-readable Markdown report, including a deterministic 95% bootstrap interval when at least three episodes exist. A training-seed overlap marks the split contaminated; dirty dataset source, dirty evaluator source, mixed tools/distributions, or fewer than three episodes makes the report non-comparable. Producer and checkpoint identity come from the dataset contract and cannot be relabeled by the report command. A clean held-out result is still reported as `screen-only`. Simulation ranking has not yet established correlation with physical rollouts, and no score from this command is motion or human-contact authorization. Run a checkpoint closed-loop through the same async LeRobot server used by a rollout: ```bash scripts/eval/serve.sh --policy /models/candidate --env-root ~/il-serve tatbot sim eval policy -- \ --server 127.0.0.1:8080 --policy /models/candidate \ --wire-scenario act_rgbd14_masked --distribution paper-draw \ --repetitions 3 --seed 20260903 \ --output-dir ~/tatbot-sim/eval/candidate-20260903 ``` A blank-sheet control runs the same worker with no server and no checkpoint: ```bash tatbot sim eval policy -- \ --client-mode hold-control --distribution paper-draw \ --repetitions 3 --seed 20260910 --output-dir ~/tatbot-sim/eval/hold-20260910 ``` It sends the worker's own joint state back every step, scores F1 0, and runs from the simulator's interpreter, so a sim node without the serving environment can produce it. Its report carries the checkpoint id `hold-control` and a digest of that contract name, never a model digest. The LeRobot client and ManiSkill worker remain separate processes and virtual environments. Their local socket carries bounded JSON plus typed arrays, while the client deliberately uses the deployed gRPC inference protocol. Feature keys and shapes come from the actual Tatbot follower declaration. Chunk overlap, stale-timestep rejection, 30 Hz target filtering, joint slew, worker protocol, checkpoint digest, zero-effort basis, surface profile, tool geometry basis/warnings, intended/drawn/overlay hashes, and optional videos are retained in the episode bundle. `--resume` reuses only completed deterministic episodes. The fixed spiral is reserved for regression controls. Normal policy batteries use the seed-generated design stream (or explicitly acquired DBV3 artwork), so a successful client cannot overfit a small checked-in design catalog. ## Contribution checklist 1. Add a deterministic fixture for the new behavior. 2. Assert units and coordinate frames at the boundary. 3. Run `scripts/check --light`. Run `scripts/check sim` where a locally qualified simulation profile is present; otherwise preserve its explicit profile-missing skip and exercise the config-independent tests you changed. 4. Label simulated results as simulated in the run manifest. ## Shared artwork and calibration Normal body, paper and skin drawing use `dbv3-acquired-v1`, the three native DBV3 acquisitions in Inkmap's manifest. The train, validation and test splits hold out their source families. Each acquisition has one physical size and pen configuration; resizing or changing pen width requires regeneration. Acquired paths retain order and direction through the simulator schedule. The spiral remains an explicit simulator calibration control, never a finished artwork fallback. `--generated-design-dir` now reads native acquisition directories with `result.json`, `artwork.json` and a matching frozen recipe. Older traced-image materializations require new DBV3 acquisition. ```bash tatbot sim sample --count 12 -- --output-dir ~/tatbot-sim/artwork-train \ --artwork-split train tatbot sim sample --count 4 -- --output-dir ~/tatbot-sim/artwork-test \ --artwork-split test --poses reclined-left-arm-supported --sites forearm tatbot sim generate paper-draw -- --out-dir ~/tatbot-sim/paper-art \ --num-episodes 4 --artwork-split train tatbot sim qualify-artwork -- --output-dir ~/tatbot-sim/artwork-review tatbot sim qualify-artwork -- --output-dir ~/tatbot-sim/artwork-review-hatch --fill-style hatch ``` `qualify-artwork` plans the collection through the same scheduled material strokes the planner schedules for a drawing; `--fill-style` selects the paint planner ([fill styles](design-format.md#fill-styles)) and the report records it, so the two planners can be compared artwork by artwork on stroke count, ideal footprint IoU and simulated deposition. ### Recipes `tatbot sim recipes` expands a frozen artwork library into reproducible scenario recipes without compiling, rendering or executing anything: ```bash tatbot sim recipes -- --output-dir ~/tatbot-sim/recipes-v1 --count 1000 --seed 7 tatbot sim recipes-status -- --output-dir ~/tatbot-sim/recipes-v1 # native DBV3 acquisitions instead of the bundled collection: tatbot sim recipes -- --output-dir ~/tatbot-sim/recipes-gen --count 1000 \ --artwork-dir ~/tatbot-artwork/dbv3-acquisitions ``` A recipe is a pure function of the frozen plan and its index. Every axis — artwork, identity, pose, site, scale, rotation, mirror, target pose, and the named camera/appearance/sensor/compile streams a later stage will use — is drawn from its own seed derived from (plan seed, recipe key, axis name). So sharding the work, resuming it, or reordering it produces byte-identical recipes, and the plan's digest is the run's identity. Rerunning the command resumes: recipes already on disk are verified against the plan and kept, anything that no longer matches is rewritten from the plan rather than trusted. The ledger counts `requested`, `admitted`, `rejected`, `compiled`, `rendered` and `executed` separately, and the last three stay zero here. A recipe is not a render, and a render is not a completed drawing. Rejected recipes keep their reason in `rejected.jsonl`. Splits are assigned from the artwork's **family** and the identity, before any augmentation — never from a variant's seed, crop or SVG digest — so a rotated or rescaled copy of a training artwork cannot become held-out artwork. The command audits family and identity leakage across the held-out partitions and exits non-zero if it finds any. Artwork with no reviewed family — a freshly generated library, for instance — is one group, so the whole of it moves across splits together. That is the conservative reading of unknown provenance rather than a meaningful family split, and the ledger names the artwork in `unknown_family_artworks` instead of leaving it to be inferred from a single-split tally. Families are reviewed metadata; nothing here invents one from a subject line. The collection binds exact acquired JSON hashes, source/license metadata, recipe identity, generation width, frozen physical size, and artwork families. The three bundled families have one training, one validation and one test example. Placement rotation and scene variation retain the family's split; duplicates cannot cross it. A different physical size requires regeneration through native DBV3. Episode metadata records source hashes and family/split. Evaluation reports block artwork from training or unknown splits and report calibration controls separately from artwork comparisons. Historical datasets are not relabeled. Normal sampling and semantic resolution compile typed v3 bundles. The placement optimizer authors new immutable bundle requests for each candidate, binding body position, yaw, and bounded support offset. Its candidate/rejection ledger lives in the suite manifest; it does not mutate an already compiled v3 world transform. All existing IK and robot/tool-shaft clearance gates remain. Paper/skin artwork defaults to a 9,000-step (five-minute) cap at 30 Hz — `config.ARTWORK_HORIZON_STEPS`, raised from 3,600 on 2026-09-10. Complete artworks must fit; randomizing an artwork never drops a stroke to meet the cap. The cap is not free: a candidate that used to be refused cheaply at two minutes now runs the full IK reach audit over three times the trajectory, so a suite that leaves the audit on takes correspondingly longer. `--no-reach-audit` skips the CPU yaw selection; final generation still enforces exact FK. `tatbot sim sample -- --max-seconds N` bounds a suite's wall clock (default 1800, `0` for none). The budget is enforced in two places because one is not enough: between candidates, and as a share of the remaining budget around each placement search. Wrapping only the last stage of that search left a measured run twenty minutes past a ten-minute budget without the check ever being reached. A search that outruns its share is refused with reason `time_budget` and recorded like any other rejection; the suite then finishes with the honest partial report it already produces for an incomplete run. The default footprint is a **nominal simulated 0.3 mm**, not measured tool calibration. V3 episodes use their bound uniform footprint width; mixed-width episodes refuse until per-stroke width execution is available. Continuous contact is sampled between control frames so a narrow line does not become disconnected dots; pen lifts and resets break that interpolation. ### What the planner will and will not draw Two geometric gates decide whether artwork becomes a drawing, and both moved on 2026-09-10 at the fleet owner's direction. Neither is a safety limit: they gate fidelity, not motion, contact force, retract or ink accounting. **Paint coverage** (`fill_geometry.PAINT_COVERAGE_FLOOR`, 0.90, was 0.98) is how much of a painted region the planned stroke path must actually cover. The old threshold refused 57 of 95 candidates in a measured suite over generated flash. Those refusals are not one population: their coverage runs 0.0, 0.31, 0.77 … 0.97, and a floor of 0.70 admits exactly the same set as a floor of 0.50, because a handful cover essentially nothing. 0.90 leaves a tenth of the ink at most missing, and a drawing that covers nothing is still refused. It costs little against a looser floor: the same library gives 9 usable artworks at 0.98, 11 at 0.90 and 12 at 0.80. A consumer that wants the old strictness passes `coverage_floor=0.98`. **Paint thinner than the tool** used to refuse the whole drawing. A tool that cannot draw a line thinner than its own tip does not refuse to draw the line: it traces the middle and the line comes out at tool width, which is what a person does with a fine liner. Those regions are now traced down the middle, and ink is allowed outside the artwork only in the halo the tool needs for them. How much heavier the result may be is bounded by `NARROW_OVERDRAW_LIMIT` (4x the painted area): a 0.3 mm tool on a 0.2 mm line lays 1.76x and is drawn; on a 0.05 mm hair it would lay far more and is refused, as is any tool larger than the region itself. Measured on a 24-piece generated library at 30-45 mm: 11 of 24 compile to strokes, against 9 before these changes. Of the twelve that do not, four carry paint finer than the tool can meaningfully trace and five cover essentially nothing at any size — properties of what the model drew. Generate roughly twice the artwork a suite needs and expect to discard about half. Qualification compares source SVG, shared typed preview, canonical target, ideal finite-width coverage, complete expert trajectory duration, and runtime InkField deposition. The reference uses 48 px/mm with 4x target supersampling; paint comparisons require IoU >=0.98 and runtime deposition requires >=0.90. The latter is a discrete nominal pigment model, separate from source-paint parity. Each size, rejection, duration, stroke count, and pen lift is recorded. Output includes a visual contact sheet and source/compiled/target/deposition images. This CPU planar qualification does not establish GPU scene, robot IK, measured deposition, or physical execution acceptance. The body sampler and bounded generated episodes provide separate evidence for their own stages. Artwork generation resolves pigment and pad textures at 16 px/mm (the flat recipe can override `--artwork-pixels-per-m`). Typed body patches use the same 16 px/mm pigment minimum; reference targets retain their independent minimum. The higher resolution costs GPU memory, so batch size must fit the selected node. Nominal width is 0.3 mm, independent of the original 2–4 mm legacy ink DR. Raster imports use the shared polygon tracer, retaining small connected marks and holes; spline fitting is excluded because it can distort paint without reporting an error. Regenerate the five-pose artwork gallery with `tatbot sim showcase-artwork -- --output-dir /path/outside/checkout`. Review and install its six JSON files in `web/inkmap/public/showcase/`; its manifest preserves offline-only qualification. The line and geometry files under `config/inkmap/examples/` remain contract fixtures, not normal design sources or gallery content. ## Dependency and check boundaries The default `tatbot_sim` install provides CPU preparation, geometry and dataset utilities. Its numerical dependencies include PyTorch; it does not install ManiSkill or SAPIEN. Use `uv sync --project python/tatbot_sim --extra maniskill` on a simulation host. Generation, preview and cinematic launchers select this extra explicitly. Policy evaluation with a separately provisioned worker interpreter requires the same extra in that interpreter. `tatbot check sim` runs the partitions below. `tatbot check sim-fast` selects the same partitions and excludes marked slow integration cases. Neither runs in the push hook. Each missing capability reports its own SKIP; it cannot skip contracts or count as full coverage. | Check | Scope | Additional requirements | | --- | --- | --- | | `sim-contracts` | Resolved configuration, observation/transport contracts, CPU imports | Base Python environment | | `sim-geometry` | Numerical geometry, planning and design compilation | Node with TypeScript support, C++ toolchain, cached robot assets | | `sim-engine` | ManiSkill adapter and material/control contracts | Engine extra, cached robot assets and compiler toolchain | | `sim-render` | Rendered sensors, shared episodes, dataset round trips | Engine extra, render device and compiler toolchain | Individual groups can run with `tatbot check sim-contracts`, for example. Slow rendering cases need the C++ planner, a render device and the checkout's retained tool calibration; missing toolchains and the public nominal profile report explicit skips. `sim-fast` excludes them. This is developer regression coverage and adds no hardware-execution gate. The public profile builds its compiler fixtures only for groups that use them. Engine assets retain ManiSkill's cache layout; reading the cache path or constructing texture files does not import the physics engine.