Project · XIAO ESP32S3 Sense · tested on hardware
Fourteen words of C for the camera and the SD card. A page of Forth for when to actually press the shutter — the part you get wrong, and the part you can now fix without unmounting anything. The trigger in this sample was redesigned at the prompt, against a live sensor, after the first design failed on real hardware.
Capture policy is the part of a trail camera you get wrong.
How long between frames. What counts as motion worth keeping. What to do when the card is nearly full. Every one of those is a number you discover by watching the thing run in the place it will live — and every one of them normally costs an edit, a rebuild, a reflash, and a camera that has to come off the tree and back to your bench.
Here each one costs a line of typing, with the camera still pointed at the scene.
Hardware: Seeed Studio XIAO ESP32S3 Sense — camera and microSD on the expansion board, 8 MB octal PSRAM, native USB console. The sensor is an OV2640 on older units and an OV5640 on newer ones; the driver probes and the pin map is the same either way. Measured on the test board: 394 KB flash, 27.8 KB internal RAM, ~21.4 KB per JPEG at 800×600, 872 µs per capture.
Mechanism in C, policy in Forth. The line between them is the whole design.
| Layer | Contents | Changes how often |
|---|---|---|
| C vocabulary trailcam_words.c | 14 words: capture a frame, write a file, read free space, dump a frame over serial, wait | Once. It is mechanism, and mechanism is not what you got wrong. |
| Forth policy policy.fs | Interval, trigger threshold, disk guard, counters, calibration helper | Constantly, at the prompt, while it runs |
cam-snap ( -- len ) capture a JPEG, hold it, push byte length cam-release ( -- ) return the held frame to the driver pool cam-quality ( n -- ) 10 (best) .. 63 (worst) cam-size ( n -- ) 5=QVGA 8=VGA 9=SVGA 10=XGA 13=UXGA cam-dims ( -- w h ) dimensions of the held frame cam-dump ( -- ) base64 the held frame to the console sd-write ( n -- flag ) write held frame to /sdcard/IMG_<n>.JPG sd-free ( -- mb ) free megabytes sd-count ( -- n ) .JPG files on the card sd-ok? ( -- flag ) is the card mounted sd-mount ( khz -- flag ) try mounting at a given SPI clock sd-info ( -- ) dump what the card reports about itself ms ( n -- ) delay us ( -- t ) microseconds since boot
cam-snap holds the frame until you write it or release it. Policy decides which — which is precisely why sd-write does not release on your behalf.
No PIR, no frame differencing, no extra silicon.
A JPEG of a static scene compresses small. Something walking into frame adds edges and texture; detail rises, and the encoded size rises with it. The encoder has already done that work by the time cam-snap returns a length — reading the number costs nothing.
That part is true. What is not true is the obvious next step.
Measure the quiet scene once, pick a floor above it, keep anything bigger. On the bench that looked fine. On the board it fell apart in minutes. Here are two measurements of the same motionless scene, a few minutes apart:
| Measurement | Range (bytes) | Spread |
|---|---|---|
| First | 21296 – 21404 | 0.5% |
| A few minutes later | 21395 – 23086 | 7.8% |
The sensor's auto-exposure and gain drift, and the baseline drifts with them. A floor set at 22000 — comfortably clear of the first band — fired on 10 out of 10 frames once the scene had drifted up underneath it. A fixed threshold is not a motion detector; it is a slowly-arming trip wire.
This is the finding worth taking away from the page. The simulator that verified the Forth before flashing modelled a stable baseline, because that is what the idea assumes. Only the sensor knew otherwise.
The baseline has to follow the scene. An integer moving average with roughly an eight-frame time constant, and a trigger set as a percentage above it:
: learn ( n -- ) base @ dup 8 / - swap 8 / + base ! ; : motion? ( n -- n flag ) dup 100 * base @ 100 margin @ + * > ;
Compared as n*100 > base*(100+margin) to stay in integers — at ~17 KB frames both sides sit near 1.7 million, well inside a 32-bit cell.
Well enough to be useful, and less well than the idea promises.
The board has no display, so the honest way to answer this was to add a fourteenth word — cam-dump, which base64s the held frame to the console — recover the actual JPEGs, and look at them.
| Scene | Frame size | Confirmed by |
|---|---|---|
| Empty — ceiling, wall, one rail | 16777 median, 0.4% spread over 20 frames | recovered image |
| Hand covering ~40% of frame | 19166 — +14.2% | recovered image |
Both real values fed through the real predicate, on the device:
ok> 16777 base ! 10 margin ! ok> 16733 motion? . drop \ empty-scene minimum 0 ok> 19166 motion? . drop \ measured hand frame -1 ok> 15 margin ! 19166 motion? . drop 0 \ correctly above the hand's +14.2%
So the usable margin window for this scene is roughly 1% to 14%. That is narrower than it looks, because the signal tracks texture, not area. A smooth, evenly-lit hand against a plain wall swaps one low-detail region for another and barely moves the encoder. A textured subject against a flat background gives far more headroom — and a flat subject against a busy background makes the frame smaller, which this one-sided test will never see.
Run sizes on your own scene with your own subject before trusting any number on this page. The technique is real; the margin is scene-specific and can be nearly zero.
The first moving average learned on every frame. With an eight-frame time constant, a subject that enters and stays becomes the new baseline in about three seconds — the camera notices the arrival and then goes blind to it. step now calls learn only in the else branch:
: step cam-snap motion? if save-frame drop \ triggered: do NOT learn else +dropped learn \ quiet: track the scene then cam-release ;
The fixed threshold was diagnosed, the moving average was written, the margin was tuned, and the whole thing was re-verified against a live sensor — at the ok> prompt, with the camera never stopping. The firmware on the board is the same image throughout.
That is the argument for this project, and it is also how the bug was found. A build-flash-observe loop would have shown the same drift eventually; it just would have taken an afternoon instead of ten minutes, and the temptation would have been to nudge the constant rather than change the design.
The same mechanism is what tutorial 10 teaches and what MagNET's hive uses to push signed role bundles to nodes it cannot reach. Here watch and guarded call step by name through the dictionary, so redefining step changes a loop that is already running.
\ stop filling a card that is nearly full
: room? ( -- flag ) sd-free 20 > ;
: guarded ( n -- )
0 do
room? if step then
150 ms
loop ;
GPIO21 is both the user LED and the SD chip select. Seeed's own pin table lists that one pin under both names. You cannot use the LED as a status light while the card is mounted — driving it low asserts chip select. It will flicker on every write, and that is normal, not a fault. This sample never touches it; for a status light, use a free pin (D0–D5).
LEAVE. The engine cannot exit a do loop early, so guarded runs its full count and simply stops writing once the card is low.+!. Increments are spelled x @ 1 + x !.us wraps about every 71 minutes — a cell is 32 bits, esp_timer_get_time() is 64. Good for timing a word, useless as a clock.host.max_freq_khz from 20 MHz to 10000.Run on a XIAO ESP32S3 Sense, 2026-09-19 — ESP32-S3 rev v0.2, OV5640, ESP-IDF 5.3.1.
| Check | Result |
|---|---|
| Engine self-test on the board | 58 passed, 0 failed — 20.3 ms |
| FFI suite | 8 passed, 0 failed, MAC matching the bootloader |
| Camera capture | ~21.4 KB at 800×600, 872 µs per frame |
policy.fs over the serial link | 0 errors, 98 dictionary entries |
| Adaptive trigger | 0 false positives in 40 live frames; fires on a step change |
| Real motion | Hand measures +14.2% against a 0.4% noise floor; predicate fires at margin 10 |
cam-dump | 19166-byte frame recovered over serial — valid JFIF 800×600 |
| Build size | 394 KB flash, 27.8 KB internal RAM |
| microSD | Mounts at 20 MHz — SD02G, SDSC, 1889 MB, sector_size=512 |
sd-write / sd-free / sd-count | Verified on the card — JPEGs land, counters track |
| Capture + write | 140–202 ms, mean ~160 ms — 180× the capture alone |
Full policy.fs pipeline | 12 forced captures written; 20 consecutive frames correctly dropped |
One caveat on the motion figure. It comes from two stills — one with a hand in frame, one without, both visually confirmed — plus the predicate evaluated on those real values. It does not come from a single continuous run that caught a hand crossing the lens: four live windows were run and none happened to contain one. The sustained-trigger behaviour of the learn-freeze fix is therefore reasoned from the measurements rather than observed directly.
Two real bugs were found on the way. ESP-IDF defaults to CONFIG_FATFS_SECTOR_4096, sized for wear-levelling on internal flash, which compiles FATFS with a fixed 4096-byte sector and makes every SD card fail to mount with FR_NO_FILESYSTEM (13) — this card reports sector_size=512, so it could never have worked. sdkconfig.defaults now sets CONFIG_FATFS_SECTOR_512=y.
And shot# restarts at zero whenever policy.fs is reloaded, so sd-write silently overwrites images already on the card. Seed it past what is there — sd-count shot# ! — before any real deployment.
Budget for the write, not the capture. A frame costs 872 µs to grab and another ~160 ms to commit to the card. That caps sustained capture near six frames per second and makes the 150 ms loop delay comparable to a single write.