- JavaScript 49.5%
- Python 28.7%
- CSS 17.8%
- Shell 3.6%
- HTML 0.4%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| dist | ||
| docs | ||
| node_modules | ||
| scripts | ||
| server | ||
| src | ||
| index.html | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
Walrus Physics Club
A small AI-physics arcade built around PolymathicAI's WALRUS continuum-dynamics model.
The premise is deliberately wrong: give a learned continuum simulator progressively less appropriate jobs, expose the failure modes, and see whether any repeatable wrong behavior becomes a mechanic.
v0.2 field-test build
This pass is aimed at making the first playtest useful instead of merely pretty.
- Seven experiments share one simulation/backend contract.
- Every cartridge now has a small field-trial objective and live progress/result state.
- Neural Pinball has an actual conventional ball on the frontend. The field is sampled as a force source while simple deterministic collision keeps the ball legible.
- Pause/resume, single-step, repeatable reset, explicit reseed, camera reset and keyboard shortcuts are wired in.
- Left drag is reserved for experiment tools; right drag controls the orbit camera.
- Procedural Web Audio feedback is optional and off by default.
- Pointer interventions have field-space ripple feedback.
- Debug metrics now include connected regions, centroid, compactness, wall count and a deliberately rough anomaly score.
- The lab reports simple phenomena such as fragmentation, mass-accounting failure and excess model authority.
- Fallback autoregressive depth now has the same semantic meaning as the real adapter: 8x advances eight complete simulation predictions.
- Wall rendering uses instancing rather than creating a separate draw object for every occupied cell.
- Server configuration uses an in-game dialog instead of a browser prompt.
The first pass includes these experiments:
- Soup: interactive continuum baseline.
- Drop Test: persistent external forcing.
- Blob Fight: pseudo-collisions between coherent regions.
- The Wall: neural dynamics around authoritative geometry.
- Is This a Box?: object permanence without objects.
- Neural Pinball: field-as-force-generator hybrid physics.
- Time Crimes: increasing autoregressive depth and model authority.
What is actually working now
The repo is runnable without the model. BACKEND_MODE=mock uses a deliberately simple NumPy field simulator with the same API and five-channel state layout as the WALRUS adapter. This lets the game, rendering, experiments and interaction design evolve independently of model setup.
When BACKEND_MODE=walrus, the server loads the real checkpoint using the same general path as PolymathicAI's non-Well inference example: six-frame history, field-index mapping, formatter, reversible normalization, delta prediction and autoregressive history.
The frontend is Three.js + Vite without a UI framework. The browser renders continuously while the server can run much more slowly.
Quick start: fallback physics
Requires Node 20+ and Python 3.11+.
./scripts/setup.sh
./scripts/dev.sh
Open:
http://localhost:5173
The simulation API listens on 8765. Both processes bind to 0.0.0.0, so on the Bosgame the UI is available from another machine at:
http://BOSGAME_IP:5173
The frontend automatically targets the same hostname on port 8765. The SERVER control can override that address and stores it in localStorage.
Controls
Left drag selected experiment tool
Right drag orbit camera
Mouse wheel zoom
1-5 select available tool
Space pause / resume
. advance one request while paused
R reset the current seed
N create a new seed
C reset camera
M toggle procedural sound
Esc close server dialog / return to menu
Walrus Trust controls how much authority the learned prediction receives. In real WALRUS mode, 0% keeps the previous authoritative state, 100% accepts the predicted state, and values above 100% deliberately extrapolate past the prediction.
Autoregressive Depth asks the backend for 1, 2, 4 or 8 complete prediction steps before returning a rendered state. At 8x, the model is consuming its own predictions repeatedly before the player gets another authoritative intervention.
Real WALRUS mode
Do the normal setup first. Then install the upstream project into the same Python environment. On Strix Halo, preserve a known-good ROCm PyTorch install instead of blindly replacing torch during dependency setup.
git clone https://github.com/PolymathicAI/walrus.git ../walrus-upstream
.venv/bin/pip install -e ../walrus-upstream
./scripts/fetch_walrus.sh
./scripts/run_walrus.sh
If the model/config already live elsewhere:
export WALRUS_CHECKPOINT=/path/to/walrus.pt
export WALRUS_CONFIG=/path/to/extended_config.yaml
BACKEND_MODE=walrus ./scripts/dev.sh
Check the environment without launching the game:
.venv/bin/python server/check_walrus.py
Useful environment variables
BACKEND_MODE=mock|walrus
GRID_SIZE=48 # clamped to 32..96
PORT=8765
WEB_PORT=5173
WALRUS_CHECKPOINT=...
WALRUS_CONFIG=...
WALRUS_DEVICE=cuda|cpu
Start at 48x48. Once real inference works, benchmark at least 32, 48 and 64 while recording single-step and 4x rollout latency.
Architecture
Three.js / browser
experiment UI + field trials
pointer interventions + visual feedback
interpolated field renderer
simple conventional pinball ball
|
| HTTP, roughly 5 Hz by default
v
FastAPI simulation server
per-browser world state
six-frame history
experiment interventions
metrics / phenomenon probes
|
+--> MockPhysicsEngine (NumPy)
|
+--> WalrusPhysicsEngine
formatter
reversible normalization
WALRUS prediction
delta -> state
autoregressive history
Tests
Backend/state tests:
PYTHONPATH=server .venv/bin/python -m unittest discover -s server/tests -v
Frontend syntax check:
find src -name '*.js' -print0 | xargs -0 -n1 node --check
Production build:
npm run build
Important caveat
The v0.2 package has been exercised against the NumPy backend and the FastAPI reset/step flow. The actual WALRUS checkpoint cannot be run in the packaging environment, so the first Bosgame WALRUS launch remains the real adapter/inference integration test. See docs/KNOWN_LIMITATIONS.md before interpreting any strange model behavior as a game mechanic.
What I would do after the first real playtest
The next work should be driven by actual WALRUS output rather than more speculative UI work:
- record real M5 inference timings at 1x/2x/4x rollout;
- capture/replay interesting seeds and intervention sequences;
- add an input/prediction/delta history scope;
- calibrate each trial around what the real model can actually do;
- replace the simple pinball collision layer with Rapier only if the field-force idea is fun enough to justify it;
- turn reproducible weird states into a persistent discoveries/materials notebook.
See docs/EXPERIMENTS.md, docs/WALRUS_NOTES.md and docs/KNOWN_LIMITATIONS.md for more detail.