Tensegra
Engineering reliable agents through programmable attention, typed symbolic execution, dependency-aware reuse, and causal evaluation: isolate model–tool failures, quantify useful computation, and test repairs against verified task utility.
I built Tensegra to make agent reliability and compute cost measurable at the level where they arise: attention routing, operand binding, dependency validity, execution, and return consumption. A useful agent must distinguish reusable work from stale work, allocate another inference or tool call only when it can change the decision, and preserve exact intermediate results through a distributed neural state. Tensegra provides the PyTorch modules, protected runtime interfaces, procedural environments, and causal diagnostics needed to engineer that boundary. The application targets are dependency-aware workflow automation, tool orchestration, and adaptive inference; the evaluation substrate is deliberately controlled.
GitHub: source, protocols, and audits · Hugging Face: analyses, checkpoints, and source snapshots
Problem
An agent's cache key is not its validity condition. A retrieved computation may match the current query while depending on an obsolete requirement, a different entity binding, or an invalidated predecessor. Similarly, a correct auxiliary prediction is useful only if the policy reads it, and an exact tool return is useful only if subsequent state updates preserve its decision-relevant information. These are distinct interface failures with different remedies. I make them separately observable through explicit entity identity, typed relations, admissibility checks, state-transition records, and interventions on the values crossing each boundary: debugging mechanisms, not just scores.
Solution
The architecture separates a continuous workspace from a symbolic store with stable identities and typed operations. Recurrent microsteps refine the workspace, propose candidate-local operations, estimate readiness, and conditionally commit admissible transitions. A return encoder maps the resulting typed event back into the next recurrent update. Schematically, , when the gate admits the proposal, and . Here denotes current public input; the notation summarizes the interface, not a claim that all experiments share one controller.
Crystallization is a local commitment: selected latent components become explicit operator and operand bindings with a validity contract. Symbolic computation advances through protected microsteps while the surrounding latent state continues to evolve; the encoded result then disperses back into the neural workspace. This separates execution authority from representation learning and makes auditable tool execution an architectural boundary rather than an interpretation of generated prose.
How
Programmable attention geometry. A runtime graph carries batched entity identities, immutable identity features, relation tensors , and padding masks independently of token order. Separate query/key projections produce soft bindings and with an edge-free null slot. After excluding that slot, each relation induces a token-space bias . Attention combines content logits with weighted structural biases, , or applies an explicit legal-read mask. Recomputing grounding from the recurrent residual allows the query's evolving identity to select a new symbolic neighborhood. Soft bias, hard locality, graph-as-input controls, and explicit pointer writes are separate experimental arms; they confer different structural information and computational authority.
Typed lowering and protected state. The recurrent execution studies factor candidate construction into operator choice, argument references, type compatibility, and readiness. Identity-only grounding, cosine-normalized projections, attention-weighted immutable-key writes, and append-only exact execution isolate binding stability from content computation. A proposal can therefore be inspected before commitment, and an execution event can be traced to its operands and preconditions. Return-drop, return-shuffle, and wrong-value interventions test whether subsequent decisions depend on the executed result; first-error, persistence, recovery, and reconvergence statistics locate where the state trajectory diverges. This is the machinery for fault isolation across model–tool boundaries.
Dependency-aware reuse and metacontrol. The later environments expose evolving requirement versions, commitments, retrieved records, and applicability relations. Policies choose computation, retrieval, validation, reuse, probing, revision, or stopping under resource charges. Verified success and accumulated cost define task utility; model-selection criteria distinguish a cheaper successful trajectory from an early unsuccessful halt. Strategy-selection worlds expose public information separately from hidden-state oracles, so a learned controller can be compared against an attainable simple-policy baseline before committing more training budget. Measure headroom before scaling is a concrete decision rule here: establish what public information can buy, identify the residual regret, then test the proposed mechanism against that residual.
Tests
The harness couples initial parameters, data schedules, and procedural populations across ablations; separates training, selection, and held-out seeds; freezes source/configuration identities; and retains process-level resource receipts. Relation removal and permutation probe structural dependence, while oracle bindings and exact traces distinguish execution capacity from policy acquisition. Candidate-local/global gates, fixed/recurrent compute, and protected/learned transitions isolate interface choices rather than changing several factors at once. Independent analysis scripts reconstruct reported endpoints directly from raw records.
For the factor-consumption experiments, a genuine shaping × reading design separates auxiliary gradients through the shared trunk from explicit factor inputs to the policy. Exact, zero, mean, and in-support substituted factors test causal consumption; out-of-fold and noise-mixed training contracts control exposure to predictor error. Counterfactual octets vary factor combinations while preserving matched instances. Flip, invariance, and near-miss populations distinguish correct action changes from indiscriminate sensitivity or conservatism. Per-seed contrasts and two-level bootstrap intervals, where specified, separate configuration uncertainty from seed variation. The result is causal evaluation for architecture decisions, with acquisition error, representation use, and interaction discrimination measured independently.
Results
The 570-run structural-attention study covers efficiency, topology corruption, graph transfer, heterogeneous mechanisms, and learned biases. The recurrent workspace study retains 780 declared evaluation rows, 11,520 training rows, and 225 checkpoints across 45 seed/arm runs. In Campaign03, full-input imitation bootstrap endpoints reached 0.93–0.98 sealed IID success with near-zero invalid reuse; applicability interventions connected behavior to the supplied dependency relations. These are bounded endpoint measurements with explicit supervision and information contracts.
The scope remains research: free recurrent policies did not acquire reliable exact execution and return reintegration, later RL degraded competent bootstrap policies, and the final consumer intervention did not pass its composition gate and replicate. The synthetic tasks supply structural/semantic interfaces; production savings and general autonomous symbolic reasoning are not established. The recurrent execution report and final mechanism audit document those boundaries.
Lessons
My engineering contribution is the decomposition from a broad reliability problem into replaceable, instrumented contracts: perception → binding → admissibility → execution → return encoding → policy consumption. That decomposition turns “the agent got stuck” into a testable distinction between a stale dependency, an incorrect operand, premature readiness, a lost return, and a policy that ignores valid information. Reproducible systems diagnosis is the transferable capability: construct the smallest environment that preserves the failure mechanism, establish strong matched controls, account for compute, and connect a repair to verified task utility.
The research record links implementations to frozen protocols, analyses, and independent audits. The Hugging Face archive holds campaign03–07 analyses and checkpoint artifacts with source snapshots and environment metadata, providing a concrete starting point for reproducing the mechanisms and testing a different controller.