RRP: relational attention for a structured latent packet

Research toward morphology-general robot control: a planner samples a structured latent packet z[K×M×D] that carries the task's meaning, a body-side controller realizes it from proprioception and touch, and every attention layer adds a sum of declared relational terms from one registry of 75 factors. Simulation only; a pre-registered campaign is running.


RRP asks one question. What must a robot policy say about its intent if the policy must survive a change of body? The test answer is a latent packet. A planner emits a continuous tensor z∈RK×M×Dz\in\mathbb{R}^{K\times M\times D}: one slot per knot time and per body assembly. A fast controller on the body turns the received packet into joint targets. It uses only what that body senses. Task meaning must live on zz. The project requires two kinds of evidence for this: probes, and causal edits of the packet.

The name began as relational robot policy. The repository is now structured-psi0-latent-diffusion-dynamics. All work is in simulation. No physical robot was commanded.

privileged teacherAn RL expert with a privileged height scan walks h1 up the step course. Reference only. Not a policy under test.
learned · R2 latent routeSystem 1 → packet → System 0. Panda pick-and-place, success. 1.5× speed.
learned · BCBehaviour-cloning pointer on a Computerworld dev seed: 8 + 5 on the calculator. Illustration, not a result.
scripted teacherPanda + UR5e bar handover. 2× speed. Dual-arm track parked.
scripted teacherFrozen tracker, ten legged and humanoid bodies. Recorded 25 Sep, before the learned legged results; the banner is from that date.

Every clip states its action source on the frame. A scripted teacher and a learned policy look the same in a video; the label is the only difference.

Problem

A policy trained on one arm usually knows nothing useful about the next arm. The interface between "what must happen" and "how this body does it" is implicit. It is buried in weights, and it does not transfer. Also, the network must rediscover structure that the engineer already knows: which joint is the parent of which, what touches what, which object the task is about, which widget a label belongs to. Often it does not rediscover it.

Solution

Two ideas. Each is separate, so each can be tested alone.

1 · The packet interface. The naming follows Ψ0\Psi_0.

  • System 1 is an encoder–decoder transformer trained by rectified flow. It reads typed context: morphology, scene, task events and roles, touch. It samples zz in 8 Euler steps.
  • System 0 runs every control tick. It builds one query per actuated joint from public morphology and the measured state. The queries cross-attend to the packet knots. The output is one normalized joint target per joint.
  • System 0 never sees the task, the instruction or an image. Meaning reaches the body only through zz.
  • Packet queries come from public morphology, not from a per-robot lookup. A new body therefore produces new queries.

2 · Relational attention as a sum of declared terms. Every attention call in the model uses one logit:

Aij=qi⊤kjd+∑f∈Bwf,h γf vf(i,j)+⟨ϕq(i),ϕk(j)⟩A_{ij}=\frac{q_i^{\top}k_j}{\sqrt d}+\sum_{f\in\mathcal B}w_{f,h}\,\gamma_f\,v_f(i,j)+\big\langle\phi_q(i),\phi_k(j)\big\rangle

Each factor ff is one entry in one registry: field × operator × form × source. It adds one term vf(i,j)v_f(i,j) in one of four forms:

  • bias — added to the logit;
  • augmentation — appended to q,kq,k, so fused attention kernels still apply;
  • gate — γf=2σ(u⋅c+b)\gamma_f=2\sigma(u\cdot c+b) scales another term by a task, goal or body summary cc;
  • mask — hard.

Relative geometry follows PaPE. Graph terms come from declared morphology and from the supplied task graph. Some pair terms cannot be read from public inputs: contact, support, held-by, handover. These are bilinear, ⟨Uxi,Vxj⟩\langle Ux_i,Vx_j\rangle. The bias is then a supervised attention subspace, and its probe estimate is usable at run time. Every coefficient starts at zero. Switching a factor on in a trained model is a no-op at step 0. A deploy guard refuses all simulator-truth sources at run time. Ground truth trains the probes. Only public or estimated values steer attention.

learned · no panel
geo.pos3d, distance part: −(r_j − r_i)ᵀM(r_j − r_i) from the cube over every pixel of a Panda pick-and-place state. Hand-set coefficients.geo.pos3d, distance part: −(r_j − r_i)ᵀM(r_j − r_i) from the cube over every pixel of a Panda pick-and-place state. Hand-set coefficients.
geo.pos3d, direction part: b̃ᵀ(r_j − r_i) from the cube, same state. Hand-set coefficients.geo.pos3d, direction part: b̃ᵀ(r_j − r_i) from the cube, same state. Hand-set coefficients.
kin.ancestor: the ten ancestors of the G1 left wrist in its declared kinematic tree. Given structure.kin.ancestor: the ten ancestors of the G1 left wrist in its declared kinematic tree. Given structure.
ix.contact: gripper–cube contact at grasp close, with the simulator contact points. Simulator label; trains the probe.ix.contact: gripper–cube contact at grasp close, with the simulator contact points. Simulator label; trains the probe.
ix.force_flow: the bottom cube of a three-cube stack carries the two above it. Simulator label closed by the flow operator.ix.force_flow: the bottom cube of a three-cube stack carries the two above it. Simulator label closed by the flow operator.
ix.held_by: the lifted cube is held by the gripper. Simulator label; trains the probe.ix.held_by: the lifted cube is held by the gripper. Simulator label; trains the probe.
task.next_contact: the posterior over the gripper's next contact after one distractor is excluded — 0.5, 0.5, 0. Synthetic evidence schedule.task.next_contact: the posterior over the gripper's next contact after one distractor is excluded — 0.5, 0.5, 0. Synthetic evidence schedule.
ui.label_for: the Name label points at its text box in a rendered Computerworld form; grey frame is ui.above. Given structure.ui.label_for: the Name label points at its text box in a rendered Computerworld form; grey frame is ui.above. Given structure.
One attention logit, one term per panel. Each panel draws one term for one query token (ring) over a real simulator state. The factor operators compute the values; they are not learned attention. Panels 1–2 use hand-set coefficients. Panels 3 and 8 are given structure. Panels 4–7 show the simulator label that trains the probe; the deployed term is the probe estimate.

How

  • Language / runtime: Python 3.12, PyTorch, one rrp CLI. A resource broker runs every heavy job on two NVIDIA GB10 machines.
  • Simulation: native MuJoCo with MuJoCo Menagerie robots; a batched Warp legged environment; the Ψ0\Psi_0 SIMPLE benchmark on Isaac Sim; Computerworld as a 3D UI world for a pointer track. 251 distinct bodies ran in at least one recorded run. 150 are in a training set.
  • Code and weights: code on GitHub; checkpoints on Hugging Face.
  • One interface each for environments, policies and tasks. One harness runs any accepted policy × environment × task. It records each declined pair with its reason.
  • Relation-aware data. One privileged state interface feeds four registries: label functions, composable scene parts, factor-agnostic transforms (reveal, surprise, counterfactual swap, noise, occlusion, subsampling) and a composition operator. A scheduler promotes each factor from isolated scenes to composed scenes as its competence rises. The scheduler is implemented and smoke-tested. There is no curriculum result.

The five environments. Arm and dual-arm frames: scripted-teacher states (privileged). Quadruped, humanoid and Computerworld frames: reset states, no controller.The five environments. Arm and dual-arm frames: scripted-teacher states (privileged). Quadruped, humanoid and Computerworld frames: reset states, no controller.

Tests

  • Merge gate: pytest tests/unit. It skips, by name, each test that needs weights or third-party assets.
  • Routes, on matched seeds: R0 scripted teacher (privileged); R1 oracle packet (diagnostic, not deployable); R2 generated packet (deployable); plus a behaviour-cloning control on the same demonstrations.
  • Edit suites change the task context and roll out.
  • Semantic and no-semantic variants have matched capacity.
  • Splits are sealed and hash-pinned before results. Each sealed cell runs once.

Results

Recorded results only, from the report snapshot of 2 October 2026.

  • Semantic supervision helps inside the training bodies. Four arm bodies, two seeds. The latent route with semantic packet supervision succeeds in 425/480 deployable episodes. The capacity-matched control succeeds in 314/480. Difference: +0.23 [0.18, 0.28], ahead in 8/8 body × seed cells. The latent route is still below plain BC on Panda: 135/180 against 58/60.
  • The packet is causally used. A "halt" added to System 1's task context can reach the body only through zz. It cuts forward progress by 0.40 m on anymal_c and 0.66 m on go2 against the control, 3/3 seeds ordered. An irrelevant context change moves the body 0.02 m at most. In Computerworld, a probe-guided packet edit moves the pointer from 163 px to 41 and 22 px of the new target. A random edit of equal norm gives 102–105 px. The edits are partial. Task success does not change.
  • Transfer to a new arm is not supported. Zero-shot on sealed xarm7 targets: latent route 0/200, BC 0/200. With target demonstrations, BC fine-tuning reaches 69–192/200. The best latent adaptation was added after the sealed result, so it is post-hoc. It beats a System 0 refit in 12/12 cells and stays below BC in 12/12. The latent route has not beaten BC on any new body.
  • Pointer. In distribution: BC 398/400, semantic latent 372, control 368. On unseen words and names, every learned method scores 0/100. None copies characters from the instruction.
  • Pretrained humanoid (Ψ0\Psi_0). Released 20/20; direct fine-tune 19/20; structured head 0/20. The cause was an integration fault: System 0 ignored the packet. This is not evidence against structure. The packet-use gate added after the fault passed. The closed-loop rerun is queued.

The report does not claim three things: that the relation factors help (the test is running), that the latent route transfers across bodies, or anything about a real robot. A ten-node pre-registered campaign is in progress: humanoids, arm diversity, Ψ0\Psi_0, the pointer and the relation factors. Its gates were written before the runs.

Lessons

The most useful part of the repository is bookkeeping. An early result said "semantic supervision hurts". It was a confound: an unbounded probe loss made System 0's updates two orders of magnitude smaller. The check found it only because every comparison has a decision number and a capacity-matched control. A source label (teacher, oracle, learned) on every frame and every table costs little. Negative results go in the same ledger as positive ones. Together these keep the project honest about where the latent route is still behind plain BC.

Every factor, one card each

75 implemented relation factors, one deck per family. Swipe a deck left or right, or use the arrows. Each card draws one logit term for one query token (the ring) over a real simulator state. The scene source is in the corner. The factor operators compute the values; they are not learned attention. For a pair term trained from simulator truth, the card shows the label that trains its probe, not the deployed estimate. Three cards are marked illustrative. Tap an image to enlarge it.

Structure

Token-to-assembly membership, slot identity, and which packet knots each System 0 node can read.

edge.node_in_assembly: a joint / sensor token attends to the assembly (arm, gripper) it belongs toedge.node_in_assembly: a joint / sensor token attends to the assembly (arm, gripper) it belongs to
edge.node_in_assembly

a joint / sensor token attends to the assembly (arm, gripper) it belongs to

source: public
1 / 9

Geometry

Relative position, depth, orientation, surface normals. geo.pos3d is the PaPE term in kernel-compatible form.

geo.above: above / below along gravity: +1 for tokens (and surface points) higher than the query by > 1 cm, -1 lower, 0 in the dead zone (query: middle cube of a stack)geo.above: above / below along gravity: +1 for tokens (and surface points) higher than the query by > 1 cm, -1 lower, 0 in the dead zone (query: middle cube of a stack)
geo.above

above / below along gravity: +1 for tokens (and surface points) higher than the query by > 1 cm, -1 lower, 0 in the dead zone (query: middle cube of a stack)

source: public
1 / 5

Kinematics

The declared kinematic tree: parents, children, ancestors, siblings, mirror pairs, limbs, and the feet and terrain cells they reach.

edge.foot_of: a limb token attends to its own foot token (feet exist only on foot-carrying assemblies; limbs drawn at hips, feet at foot sites)edge.foot_of: a limb token attends to its own foot token (feet exist only on foot-carrying assemblies; limbs drawn at hips, feet at foot sites)
edge.foot_of

a limb token attends to its own foot token (feet exist only on foot-carrying assemblies; limbs drawn at hips, feet at foot sites)

source: public
1 / 8

Interaction

Contact, support, force flow, holding, handover. Public inputs cannot give these pair terms. They are bilinear, and each has its own probe.

ix.contact: two entities in geometric contact (sim contact truth): query gripper → cube 1, distractors / target 0ix.contact: two entities in geometric contact (sim contact truth): query gripper → cube 1, distractors / target 0
ix.contact

two entities in geometric contact (sim contact truth): query gripper → cube 1, distractors / target 0

source: gt train-only + est-probe
1 / 5

Task

Edges from the supplied task graph — actors, patients, destinations, roles, dependencies, receipts — and the task-gated next contact.

edge.actor_of: manipulator entity ↔ every event it is (cooperating) actor ofedge.actor_of: manipulator entity ↔ every event it is (cooperating) actor of
edge.actor_of

manipulator entity ↔ every event it is (cooperating) actor of

source: public
1 / 15

Locomotion

The stability margin of the centre of mass, and each swinging foot's next foothold.

leg.com_support: readout of the stability margin: signed planar distance of the COM projection to the stance-foot support polygon (m, + inside)leg.com_support: readout of the stability margin: signed planar distance of the COM projection to the stance-foot support polygon (m, + inside)
leg.com_support

readout of the stability margin: signed planar distance of the COM projection to the stance-foot support polygon (m, + inside)

signed dist(COM, hull(stance feet))
source: gt train-only
1 / 2

UI

The Computerworld widget tree: same window, render order, tab order, labels, drag targets.

ui.above: signed render order: +1 if widget j is drawn above i (higher dense z-layer), -1 below, 0 same layerui.above: signed render order: +1 if widget j is drawn above i (higher dense z-layer), -1 below, 0 same layer
ui.above

signed render order: +1 if widget j is drawn above i (higher dense z-layer), -1 below, 0 same layer

source: public
1 / 5

Probes

Readouts in the same registry: what the packet is supervised to carry, read through opaque handles only. Probes are diagnostics. Edits and rollouts show causal use.

probe.arm.acting_on: readout of whether assembly a's hand links touch entity e (contact truth)probe.arm.acting_on: readout of whether assembly a's hand links touch entity e (contact truth)
probe.arm.acting_on

readout of whether assembly a's hand links touch entity e (contact truth)

source: gt train-only
1 / 26

Related

RRP: relational attention for a…multigraph-nnmultigraph-nnComputerworldComputerworldRecursive Omnimodal Video Action ModelRecursive Omnimodal Video A…SynthUXSynthUXComputatrumComputatrumTeaching Computers to Use ComputersTeaching Computers to…The Tensor ComputerThe Tensor ComputerThe Cortical CanvasThe Cortical CanvasLooped Attention in Video Diffusion TransformersLooped Attention in V…The Node Neural Network (NNN)The Node Neural Netwo…multi-graph-formermulti-graph-formerLLMs are the Update Rules of Intelligent Fractals: Escaping the Context Window with Iterative, Structured Local UpdatesLLMs are the Update…Software Engineering After AgentsSoftware Engineerin…belief-graph-orchestratorbelief-graph-orches…notion-vibestartupnotion-vibestartupNode TreeNode Treebrowser-osbrowser-oswindows-web-nextwindows-web-nextmacos-web-nextmacos-web-nextgeneral-unified-world-modelinggeneral-unified-wor…TensegraTensegraTensorCodeTensorCodeFull Stack Artificial IntelligenceFull Stack Artifici…A Differentiable von Neumann ComputerA Differentiable vo…yt2ctxyt2ctxBroadening and building beyond classical reinforcement learningBroadening and buil…Canvas Engineering: Declared Causal Macrostructure for Reverse-Diffusion Latent DynamicsCanvas Engineering:…AI systems engineeringAI systems engineer…Research and technical writingResearch and techni…👩🏽‍🌾 The Fertile Crescent👩🏽‍🌾 The Fertile…ComputatrumComputatrumDRAG TO ORBIT · SCROLL OR PINCH TO ZOOM