diff --git a/.RnD/openDerivation/godbolt/README.md b/.RnD/openDerivation/godbolt/README.md index 206b67f..1185434 100644 --- a/.RnD/openDerivation/godbolt/README.md +++ b/.RnD/openDerivation/godbolt/README.md @@ -123,13 +123,16 @@ What this shows, and what it does not: ```c++ const u8 x=inv(u8(Src::get(in))); - if constexpr (s>=0) return u8((((x<=0) return u8(u8(u8(u8(x<>(-s))+p)&m)+Base::proc(in)); ``` - the HAPI cell is 27 instructions, byte for byte `wave4_c.c`. It gives the same result as the current `proc` for every - `n`, `s` in −7..7, `p`, `m` and input byte (host check, 3.0e9 cases). This is not applied here: it changes - `examples/static_net/include/waveCell.h`, and with it static_net's measured figures. + the HAPI cell is 27 instructions, byte for byte `wave4_c.c`. It gives the same result as the earlier `proc` for every + `n`, `s` in −7..7, `p`, `m` and input byte (host check, 3.0e9 cases). Left shifts keep the early `u8`: without it + gcc 7.3 does them in 16 bits (static_net's `mixed_avr_size` grows by 30 B). This is now what + `examples/static_net/include/waveCell.h` does; static_net's `wave4` went from 44 B / 25 cycles to 42 B / 24 cycles, and + every other program of its `check/build.sh` kept its size. The sources here stay pinned to `3b0c466` and keep the + earlier spelling. - **A current gcc folds it anyway.** With AVR gcc 16.1 on Compiler Explorer, the HAPI cell and `wave4_c.c` compile to the same 24 instructions, with no `mov` and a shorter readout (`sbrc`/`inc` for `x[3]`, `cpi`/`sbc`/`neg` for the threshold). The one-instruction gap is specific to gcc 7.3, the toolchain static_net measures with. @@ -172,8 +175,9 @@ To compare two of them, open one source pane per file, each with its own compile - `wave4_apiof.cpp` and `wave4_od.cpp` give the same `wave4()`, 28 instructions, 56 B (the translator check). - That `wave4()` is instruction for instruction the `bnc_predict` static_net builds from `include/` for `compare_emlearn`. - The 44 B figure static_net reports is this function minus the null model's `bnc_predict` (12 B, a single compare), the floor - `compare_emlearn/run.sh` subtracts: 56 − 12 = 44. For plain C the same subtraction gives 54 − 12 = 42. + The 44 B figure static_net reported at that commit is this function minus the null model's `bnc_predict` (12 B, a single + compare), the floor `compare_emlearn/run.sh` subtracts: 56 − 12 = 44. For plain C the same subtraction gives + 54 − 12 = 42, which is static_net's figure since `WaveOf` was respelled. - `wave4_c.c` compiles as C++ and as C to the same 27 instructions, 54 B, and differs from the HAPI cell by the instructions listed above. - `wave4_c.c` and the HAPI cell give the same answer on all 2^32 inputs. diff --git a/README.md b/README.md index 5ce77de..5f7ad6c 100644 --- a/README.md +++ b/README.md @@ -601,8 +601,8 @@ confirmed by eye. [`examples/static_net`](examples/static_net) is a different kind of composition: **static networks**, dry, typed, zero-runtime descriptions of dataflow nets (a typelist of parts, wired by index, id or query; -the values a net reads are the slots of a state, `hapi/slots.h`, and a register is a layer of it), with tinyML as the running example: a 4-input classifier in 44 B and -25 cycles on an 8-bit ATmega328p. It is measured against a table loop and against emlearn on the same +the values a net reads are the slots of a state, `hapi/slots.h`, and a register is a layer of it), with tinyML as the running example: a 4-input classifier in 42 B and +24 cycles on an 8-bit ATmega328p. It is measured against a table loop and against emlearn on the same folds, in one build setup; the cycle counts come from simavr and were confirmed on a real Arduino Nano (100 numbers, no difference), and the realization of a cell (unrolled or as a table and a loop) is a compile-time choice with a measured size/speed trade. Its README says what was measured and what was not. diff --git a/examples/static_net/README.md b/examples/static_net/README.md index 62fa9b5..3eac56c 100644 --- a/examples/static_net/README.md +++ b/examples/static_net/README.md @@ -4,7 +4,7 @@ Static networks: dry, typed descriptions of dataflow nets. A **net** is a typelist of parts, wired by index, by id or by a query. It carries no data of its own: the values it reads live in a **state** of tag-addressed slots (`hapi::Slot`, see HAPI's README), and the net is evaluated over that state. Because the whole structure is a type, the compiler resolves it at compile time: the forms below compile to the same disassembly as the hand-written equivalents (*What the composition costs*). -The running example is a 4-input classifier on an 8-bit AVR: 44 B and 25 cycles on an ATmega328p. +The running example is a 4-input classifier on an 8-bit AVR: 42 B and 24 cycles on an ATmega328p. Parts agree on a small **contract** (a static, pure `proc(in)`); they do not inherit from a framework, and HAPI's `Chain<>` / `APIOf<>` / `Expand<>` do the composing. Neural-net cells are the running example, not the point: any static dataflow of pure stages fits (a filter chain, a control loop, a sensor fusion), and other contracts can sit @@ -19,7 +19,7 @@ using BanknoteNet = wave::Cell, wave::Wave>; BanknoteState features = banknote({variance, skewness, curtosis, entropy}); // the state: four input bytes, the only runtime data -bool y = BanknoteNet::proc(features); // a few shifts, masks and adds; 44 B of flash, 25 cycles on an ATmega328p +bool y = BanknoteNet::proc(features); // a few shifts, masks and adds; 42 B of flash, 24 cycles on an ATmega328p ``` Not the same result as [`ml_interpreter_cost`](../ml_interpreter_cost): that example is *dispatch removal* (a layered net on ARM without TFLite-Micro's interpreter), which a plain template network @@ -78,7 +78,7 @@ One build setup, one harness, flash and RAM net of a no-model program, cycles pe | model | flash (B) | RAM (B) | cycles | accuracy fold 1 / 5-fold mean (%) | |---|---|---|---|---| -| **`wave4`** (this example's cell) | **44** | **0** | **25** | 100.00 / 99.27 | +| **`wave4`** (this example's cell) | **42** | **0** | **24** | 100.00 / 99.27 | | `lin4` (quantized perceptron, int8 weights) | 112 | 0 | 62 | 99.64 / 99.13 | | `table4` (same weights, a loop over a PROGMEM table) | 86 | 2 | 127 | 99.64 / 99.13 | | emlearn tree, depth 2, uint8 features | 232 | 10 | 177 | 92.70 / 90.67 | @@ -87,8 +87,8 @@ One build setup, one harness, flash and RAM net of a no-model program, cycles pe | emlearn MLP 4-4-1 (`eml_net`, float) | 4262 | 180 | 12562 | 99.64 / 99.49 | | emlearn MLP 4-16-1 | 4550 | 564 | 33961 | 99.64 / 99.71 | -At a size near `wave4` there is no emlearn model (the smallest is 5x larger, at 90.7% mean accuracy); trees need depth 8 to reach 98.5%, at about 11x the flash and 7.6x the cycles of `wave4` for 0.7 pp -less accuracy; the MLP is the only model more accurate (+0.15 to +0.44 pp at 4-16 hidden units), at about 97x the flash and 500x the cycles for the 4-hidden-unit net, more for larger ones (soft float on a chip without an FPU). +At a size near `wave4` there is no emlearn model (the smallest is 5.5x larger, at 90.7% mean accuracy); trees need depth 8 to reach 98.5%, at about 11x the flash and 7.9x the cycles of `wave4` for 0.7 pp +less accuracy; the MLP is the only model more accurate (+0.15 to +0.44 pp at 4-16 hidden units), at about 100x the flash and 520x the cycles for the 4-hidden-unit net, more for larger ones (soft float on a chip without an FPU). Cycles of the cells are constant; the trees vary with the path (`compare_emlearn/results.md` has min / mean / max for every model and all depths). **Read with care:** Banknote is nearly separable and says nothing about a harder task; one dataset; emlearn 0.23.2 as documented (the `dtype='uint8_t'` option for the best-case trees, its float MLP with the 1/255 input scaling folded into layer 0, best of 5 restarts); emlearn's fixed-point MLP is an unfinished feature upstream and its MLP `inline` method silently emits the loadable code, so no fixed-point MLP was measured (`compare_emlearn/emlearn_repro.py`). diff --git a/examples/static_net/check/build.sh b/examples/static_net/check/build.sh index ed74cd5..0c0c55b 100755 --- a/examples/static_net/check/build.sh +++ b/examples/static_net/check/build.sh @@ -84,7 +84,7 @@ avr sugar_roll_avr_off sugar_roll_avr.cpp "1006 B / 62 B" -DSUGAR_ROLL_AT= avr sugar_unroll_hand_avr sugar_unroll_hand_avr.cpp "1006 B / 62 B"; same "SUGAR_ROLL_AT huge == the unrolled cell by hand" sugar_roll_avr_off sugar_unroll_hand_avr avr registers_avr registers_avr.cpp "220 B / 4 B" avr registers_avr_flat registers_avr.cpp "220 B / 4 B" -DFLAT # the hand-indexed twin that reads before it writes: same size -avr banknote_avr_check banknote_avr_check.cpp "2346 B / 26 B" # the Banknote cell on a board: it reports over UART +avr banknote_avr_check banknote_avr_check.cpp "2344 B / 26 B" # the Banknote cell on a board: it reports over UART if have simavr && have python3; then echo "== AVR: the same rows through the same function, simulated ATmega328p against the host, row by row" avr-g++ -std=c++17 -Os -mmcu=atmega328p -DSIM_ONCE "${INC[@]}" sonar_rows_avr.cpp -o "$W/rows.elf" 2>"$W/err" || bad sonar_rows_avr "$(grep -m1 error "$W/err")" diff --git a/examples/static_net/compare_emlearn/results.md b/examples/static_net/compare_emlearn/results.md index b8f72ac..040468b 100644 --- a/examples/static_net/compare_emlearn/results.md +++ b/examples/static_net/compare_emlearn/results.md @@ -1,7 +1,7 @@ | model | flash net (B) | RAM net (B) | on-device ok | cycles min / mean / max (net of null) | accuracy fold 1 / 5-fold mean (%) | |---|---|---|---|---|---| | null | 2342 (total) | 34 (total) | 28/274 | 19 / 19.0 / 19 | - | -| wave4 | 44 | 0 | 274/274 | 25 / 25.0 / 25 | 100.00 / 99.27 | +| wave4 | 42 | 0 | 274/274 | 24 / 24.0 / 24 | 100.00 / 99.27 | | lin4 | 112 | 0 | 273/274 | 62 / 62.0 / 62 | 99.64 / 99.13 | | table4 | 86 | 2 | 273/274 | 127 / 127.0 / 127 | 99.64 / 99.13 | | eml_mlp_h2 | 4214 | 132 | 270/274 | 7978 / 8711.6 / 9145 | 98.54 / 97.89 | diff --git a/examples/static_net/include/waveCell.h b/examples/static_net/include/waveCell.h index 2189eb9..e1e9799 100644 --- a/examples/static_net/include/waveCell.h +++ b/examples/static_net/include/waveCell.h @@ -24,9 +24,14 @@ namespace wave { struct WaveOf {template struct Part:O { using Base=O; using Base::Base; static constexpr u8 inv(u8 x) {return n?u8(~x):x;} - static constexpr u8 shf(u8 x) {return s>=0?u8(x<>(-s));} - template SNET_INLINE static constexpr u8 proc(const I& in) - {return u8(u8(u8(shf(inv(u8(Src::get(in))))+p)&m)+Base::proc(in));} + // avr-gcc 7.3 -Os (measured): a right shift is done in 8 bits only when it reads a variable and nothing narrows it before + // the mask; read as a call inside the expression, or cast to u8 first, it costs a register copy (.RnD/openDerivation/godbolt/). + // A left shift is the other way round: without the early u8 it widens to 16 bits (mixed_avr_size: +30 B) + template SNET_INLINE static constexpr u8 proc(const I& in) { + const u8 x=inv(u8(Src::get(in))); + if constexpr (s>=0) return u8(u8(u8(u8(x<>(-s))+p)&m)+Base::proc(in)); + } };}; template using Wave=WaveOf; // Src: a Field or Elem of the state template using RefWave=WaveOf,n,s,p,m>; diff --git a/examples/static_net/measure/silicon_results.md b/examples/static_net/measure/silicon_results.md index 4e346bc..b7b73aa 100644 --- a/examples/static_net/measure/silicon_results.md +++ b/examples/static_net/measure/silicon_results.md @@ -8,6 +8,10 @@ count here is a count on the chip. sum of cycles over 274 rows). The timed code has no interrupts, no caches and no wait states on an ATmega328p, so exact agreement is what a cycle-accurate simulator should give; this confirms it on the chip rather than assuming it. +These are the ELFs of that date. `waveCell.h`'s `WaveOf` has since been respelled (one register copy less on avr-gcc 7.3): the `compare_emlearn` +`wave4` program now runs 43 cycles per call in simavr (`wave4.min`/`max`, 24 net of the null program) and `measure/`'s `bn_wave` is unchanged at 23; +the new ELF has not been run on the chip. + | target | number | simavr | silicon | diff | |---|---|---|---|---| | bn_wave | wave4 | 23 | 23 | +0 |