docs: add CLAUDE.md for future Claude Code sessions
This commit is contained in:
179
CLAUDE.md
Normal file
179
CLAUDE.md
Normal file
@@ -0,0 +1,179 @@
|
||||
# Experiment_ReinforcementLearning — cursor-following lunar lander
|
||||
|
||||
A cursor-following lunar lander demo: user moves the mouse, a PPO-controlled
|
||||
ship chases it against gravity. Deliberate strict separation between game
|
||||
logic, backend, frontend, and training.
|
||||
|
||||
Live repo: `https://gitea.tes.gd/admin/Experiment_ReinforcementLearning`
|
||||
|
||||
## Architecture
|
||||
|
||||
Four processes with narrow, stable interfaces:
|
||||
|
||||
```
|
||||
[React UI] ──WebSocket──▶ [ASP.NET backend] ──stdin/stdout──▶ [Game CLI (.NET)]
|
||||
▲
|
||||
│ same protocol
|
||||
│
|
||||
[Python PPO trainer] ──┘
|
||||
```
|
||||
|
||||
- **`GameCli/`** — C# .NET 8 console app. Stateless step-driven executable.
|
||||
Reads one action per line on stdin, writes one observation JSON per line on
|
||||
stdout. Owns physics, reward, observation. No networking, no UI knowledge.
|
||||
- **`Backend/`** — ASP.NET Core Web API. WebSocket at `/ws/game`. Per
|
||||
connection: spawns one `GameCli` subprocess, runs a 50 Hz tick loop that
|
||||
calls the singleton `PolicyRunner` (ONNX) for each action.
|
||||
- **`Frontend/`** — Vite + React + TypeScript. Opens a WebSocket via a
|
||||
same-origin proxy (`/ws` → `localhost:5100`), sends `mousemove` as world
|
||||
coords, renders the ship on a `<canvas>`.
|
||||
- **`Training/`** — Python. `LanderCliEnv` wraps the **same** `GameCli` binary
|
||||
as a `gym.Env`, so training physics and inference physics are identical by
|
||||
construction. Produces `models/ppo_lander.onnx`.
|
||||
|
||||
Design spec: `docs/superpowers/specs/2026-07-17-cursor-following-lander-design.md`
|
||||
Implementation plans: `docs/superpowers/plans/2026-07-17-{01-04}-*.md`
|
||||
|
||||
## Build & test
|
||||
|
||||
Everything runs from the repo root.
|
||||
|
||||
```bash
|
||||
# Game CLI (25 xUnit tests)
|
||||
dotnet test --filter FullyQualifiedName~GameCli.Tests
|
||||
|
||||
# Backend (8 xUnit tests — includes end-to-end WebSocket smoke)
|
||||
dotnet publish GameCli/GameCli.csproj -c Release -o publish/GameCli
|
||||
dotnet test --filter FullyQualifiedName~Backend.Tests
|
||||
|
||||
# Training (11 pytest tests)
|
||||
source .venv/bin/activate
|
||||
cd Training && pytest && cd ..
|
||||
|
||||
# Frontend
|
||||
cd Frontend && npx tsc --noEmit && npm run build && cd ..
|
||||
```
|
||||
|
||||
## Run the demo locally
|
||||
|
||||
```bash
|
||||
# 1. Publish CLI (idempotent)
|
||||
dotnet publish GameCli/GameCli.csproj -c Release -o publish/GameCli
|
||||
|
||||
# 2. Start backend (loads models/ppo_lander.onnx at startup)
|
||||
ASPNETCORE_URLS=http://localhost:5100 \
|
||||
dotnet run --project Backend --no-launch-profile
|
||||
|
||||
# 3. In another shell, start Vite dev server
|
||||
cd Frontend && npm run dev # → http://localhost:5173/
|
||||
```
|
||||
|
||||
Open `http://localhost:5173/`. Status bar top-left shows connection state and
|
||||
step counter; blue crosshair follows your mouse; white triangle is the ship.
|
||||
|
||||
## Wire protocol reference
|
||||
|
||||
**CLI stdio** (line-delimited JSON):
|
||||
|
||||
- CLI startup: `{"init":{"world":[1,1],"dt":0.02,"obs_dim":7,"n_actions":4}}`
|
||||
- Step in: `{"action":<0-3>,"target":[x,y]}`
|
||||
- Reset in: `{"cmd":"reset","seed":<int>?}`
|
||||
- Step out: `{"obs":[7 floats],"state":{x,y,angle,engine},"reward":<f>,"done":<b>,"step":<i>}`
|
||||
- Actions: `0`=noop, `1`=left thruster, `2`=main engine, `3`=right thruster.
|
||||
- Observation: `[dx, dy, vx, vy, sin(angle), cos(angle), angular_velocity]`.
|
||||
|
||||
**WebSocket** (`ws://localhost:5100/ws/game` or same-origin via Vite proxy):
|
||||
|
||||
- Client → server: `{"type":"cursor","x":<0-1>,"y":<0-1>}`
|
||||
- Server → client (on connect): `{"type":"init","world":[1,1]}`
|
||||
- Server → client (each 20 ms tick):
|
||||
`{"type":"state","x":..,"y":..,"angle":..,"engine":..,"target":[..,..],"step":..}`
|
||||
|
||||
Both are mirrored in code (`GameCli/Protocol.cs`, `Backend/GameProcess.cs`,
|
||||
`Backend/WireMessages.cs`, `Frontend/src/protocol.ts`).
|
||||
|
||||
## Coordinate & physics conventions
|
||||
|
||||
- World is `[0, 1] × [0, 1]`. Y increases downward (screen-native).
|
||||
- `angle = 0` = ship pointing up. Positive angle rotates clockwise.
|
||||
- No walls, no ground. Ship can drift off-screen; only the reward pulls it back.
|
||||
- Fixed `dt = 0.02` (50 Hz). Explicit Euler integration.
|
||||
- Physics constants live in `GameCli/Physics.cs`; reward constants in
|
||||
`GameCli/Reward.cs`. All `public const`, easy to tune.
|
||||
|
||||
## PPO training
|
||||
|
||||
```bash
|
||||
source .venv/bin/activate
|
||||
|
||||
# Smoke run (~30 s):
|
||||
python Training/train.py --steps 10000 --n-envs 2 --use-dummy --max-episode-steps 200
|
||||
|
||||
# Real training (~hours on CPU):
|
||||
python Training/train.py --steps 2000000 --n-envs 8
|
||||
|
||||
# Export to ONNX (backend reads this file):
|
||||
python Training/export_onnx.py \
|
||||
--checkpoint checkpoints/ppo_lander_final.zip \
|
||||
--out models/ppo_lander.onnx
|
||||
```
|
||||
|
||||
Backend loads `models/ppo_lander.onnx` at startup and fails fast if it's
|
||||
missing. Restart the backend after re-exporting to pick up the new weights.
|
||||
|
||||
### Current state of the model
|
||||
|
||||
`models/ppo_lander.onnx` was trained for only 10k PPO steps as a
|
||||
pipeline smoke — essentially random with slight bias. The ship falls off the
|
||||
world in a few seconds under this policy. To see actual cursor-following
|
||||
behavior, run the 2M-step training pass above (see PPO training section).
|
||||
|
||||
## Config paths & env vars
|
||||
|
||||
Backend resolves paths at startup by walking up from `AppContext.BaseDirectory`
|
||||
to find `GameCli.sln` (the repo-root marker), then joining a relative config
|
||||
path.
|
||||
|
||||
- `Lander:CliPath` (default `publish/GameCli/GameCli`)
|
||||
- `Lander:ModelPath` (default `models/ppo_lander.onnx`)
|
||||
- Env-var overrides: `LANDER__CliPath`, `LANDER__ModelPath` (ASP.NET
|
||||
double-underscore convention).
|
||||
|
||||
## Repo conventions
|
||||
|
||||
- Every commit lands on `main` (no PR flow — greenfield project).
|
||||
- Author: `meelstorm <meelstorm@gmail.com>`. Use
|
||||
`-c user.email=meelstorm@gmail.com -c user.name=meelstorm` when committing.
|
||||
- Conventional commit prefixes: `feat(gamecli):`, `feat(backend):`,
|
||||
`feat(training):`, `feat(frontend):`, `fix(...)`, `test(...)`, `chore:`.
|
||||
- `.gitignore` covers `bin/`, `obj/`, `node_modules/`, `dist/`, `publish/`,
|
||||
`.venv/`, `checkpoints/`, `models/`, `tensorboard/`, `*.onnx`.
|
||||
- The trained model is NOT in git — it's a build artifact.
|
||||
|
||||
## Known limitations & future work
|
||||
|
||||
- **Model is undertrained** — see "Current state of the model" above.
|
||||
- **No walls** — a badly-trained policy loses to gravity within a second and
|
||||
the ship drifts off-screen forever. Design choice: reward should be enough
|
||||
to keep a well-trained policy in-world.
|
||||
- **Angle wrap in reward is applied** but not in `Physics.Step` — the ship's
|
||||
`angle` field grows unbounded across many revolutions. Observation uses
|
||||
sin/cos so it's immune; reward wraps to `[-π, π]` internally.
|
||||
- **StrictMode WebSocket double-mount** in dev — React's StrictMode mounts
|
||||
the `useGameSocket` effect twice, so you'll see one benign
|
||||
"closed before established" WebSocket warning per page load.
|
||||
- **Single-player only** — one CLI subprocess per connection; multi-user
|
||||
demo is fine at demo scale but there's no shared world.
|
||||
|
||||
## Debugging tips
|
||||
|
||||
- **Backend can't find CLI or ONNX at startup** — the repo-root walker looks
|
||||
for `GameCli.sln`. Both `.sln` and `.slnx` formats work.
|
||||
- **`dotnet test` in `Backend.Tests` uses the CLI** — publish first
|
||||
(`dotnet publish GameCli -c Release -o publish/GameCli`) so
|
||||
`EndToEndSmokeTests` can find the binary.
|
||||
- **`torch.onnx.export` on torch 2.13+** — pass `dynamo=False` to force the
|
||||
legacy TorchScript exporter unless `onnxscript` is installed.
|
||||
- **Vite WS proxy failing** — verify with a Node script bypassing the proxy
|
||||
(`ws://localhost:5100/ws/game` direct). If direct works but proxy doesn't,
|
||||
check `Frontend/vite.config.ts` proxy block.
|
||||
Reference in New Issue
Block a user