Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 10 additions & 6 deletions .RnD/openDerivation/godbolt/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,13 +123,16 @@ What this shows, and what it does not:

```c++
const u8 x=inv(u8(Src::get(in)));
if constexpr (s>=0) return u8((((x<<s)+p)&m)+Base::proc(in));
if constexpr (s>=0) return u8(u8(u8(u8(x<<s)+p)&m)+Base::proc(in));
else return u8((((x>>(-s))+p)&m)+Base::proc(in));
```

the HAPI cell is 27 instructions, byte for byte `wave4_c.c`. It gives the same result as the current `proc` for every
`n`, `s` in −7..7, `p`, `m` and input byte (host check, 3.0e9 cases). This is not applied here: it changes
`examples/static_net/include/waveCell.h`, and with it static_net's measured figures.
the HAPI cell is 27 instructions, byte for byte `wave4_c.c`. It gives the same result as the earlier `proc` for every
`n`, `s` in −7..7, `p`, `m` and input byte (host check, 3.0e9 cases). Left shifts keep the early `u8`: without it
gcc 7.3 does them in 16 bits (static_net's `mixed_avr_size` grows by 30 B). This is now what
`examples/static_net/include/waveCell.h` does; static_net's `wave4` went from 44 B / 25 cycles to 42 B / 24 cycles, and
every other program of its `check/build.sh` kept its size. The sources here stay pinned to `3b0c466` and keep the
earlier spelling.
- **A current gcc folds it anyway.** With AVR gcc 16.1 on Compiler Explorer, the HAPI cell and `wave4_c.c` compile to
the same 24 instructions, with no `mov` and a shorter readout (`sbrc`/`inc` for `x[3]`, `cpi`/`sbc`/`neg` for the
threshold). The one-instruction gap is specific to gcc 7.3, the toolchain static_net measures with.
Expand Down Expand Up @@ -172,8 +175,9 @@ To compare two of them, open one source pane per file, each with its own compile

- `wave4_apiof.cpp` and `wave4_od.cpp` give the same `wave4()`, 28 instructions, 56 B (the translator check).
- That `wave4()` is instruction for instruction the `bnc_predict` static_net builds from `include/` for `compare_emlearn`.
The 44 B figure static_net reports is this function minus the null model's `bnc_predict` (12 B, a single compare), the floor
`compare_emlearn/run.sh` subtracts: 56 − 12 = 44. For plain C the same subtraction gives 54 − 12 = 42.
The 44 B figure static_net reported at that commit is this function minus the null model's `bnc_predict` (12 B, a single
compare), the floor `compare_emlearn/run.sh` subtracts: 56 − 12 = 44. For plain C the same subtraction gives
54 − 12 = 42, which is static_net's figure since `WaveOf` was respelled.
- `wave4_c.c` compiles as C++ and as C to the same 27 instructions, 54 B, and differs from the HAPI cell by the instructions
listed above.
- `wave4_c.c` and the HAPI cell give the same answer on all 2^32 inputs.
Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -601,8 +601,8 @@ confirmed by eye.

[`examples/static_net`](examples/static_net) is a different kind of composition: **static networks**,
dry, typed, zero-runtime descriptions of dataflow nets (a typelist of parts, wired by index, id or query;
the values a net reads are the slots of a state, `hapi/slots.h`, and a register is a layer of it), with tinyML as the running example: a 4-input classifier in 44 B and
25 cycles on an 8-bit ATmega328p. It is measured against a table loop and against emlearn on the same
the values a net reads are the slots of a state, `hapi/slots.h`, and a register is a layer of it), with tinyML as the running example: a 4-input classifier in 42 B and
24 cycles on an 8-bit ATmega328p. It is measured against a table loop and against emlearn on the same
folds, in one build setup; the cycle counts come from simavr and were confirmed on a real Arduino Nano
(100 numbers, no difference), and the realization of a cell (unrolled or as a table and a loop) is a
compile-time choice with a measured size/speed trade. Its README says what was measured and what was not.
Expand Down
10 changes: 5 additions & 5 deletions examples/static_net/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ Static networks: dry, typed descriptions of dataflow nets.

A **net** is a typelist of parts, wired by index, by id or by a query. It carries no data of its own: the values it reads live in a **state** of tag-addressed slots (`hapi::Slot`, see HAPI's README), and the net is evaluated over that state.
Because the whole structure is a type, the compiler resolves it at compile time: the forms below compile to the same disassembly as the hand-written equivalents (*What the composition costs*).
The running example is a 4-input classifier on an 8-bit AVR: 44 B and 25 cycles on an ATmega328p.
The running example is a 4-input classifier on an 8-bit AVR: 42 B and 24 cycles on an ATmega328p.

Parts agree on a small **contract** (a static, pure `proc(in)`); they do not inherit from a framework, and HAPI's `Chain<>` / `APIOf<>` / `Expand<>` do the
composing. Neural-net cells are the running example, not the point: any static dataflow of pure stages fits (a filter chain, a control loop, a sensor fusion), and other contracts can sit
Expand All @@ -19,7 +19,7 @@ using BanknoteNet = wave::Cell<WAVE_K,
wave::Wave<Curtosis,WAVE_CURTOSIS_N,WAVE_CURTOSIS_S,WAVE_CURTOSIS_P,WAVE_CURTOSIS_M>, wave::Wave<Entropy,WAVE_ENTROPY_N,WAVE_ENTROPY_S,WAVE_ENTROPY_P,WAVE_ENTROPY_M>>;

BanknoteState features = banknote({variance, skewness, curtosis, entropy}); // the state: four input bytes, the only runtime data
bool y = BanknoteNet::proc(features); // a few shifts, masks and adds; 44 B of flash, 25 cycles on an ATmega328p
bool y = BanknoteNet::proc(features); // a few shifts, masks and adds; 42 B of flash, 24 cycles on an ATmega328p
```

Not the same result as [`ml_interpreter_cost`](../ml_interpreter_cost): that example is *dispatch removal* (a layered net on ARM without TFLite-Micro's interpreter), which a plain template network
Expand Down Expand Up @@ -78,7 +78,7 @@ One build setup, one harness, flash and RAM net of a no-model program, cycles pe

| model | flash (B) | RAM (B) | cycles | accuracy fold 1 / 5-fold mean (%) |
|---|---|---|---|---|
| **`wave4`** (this example's cell) | **44** | **0** | **25** | 100.00 / 99.27 |
| **`wave4`** (this example's cell) | **42** | **0** | **24** | 100.00 / 99.27 |
| `lin4` (quantized perceptron, int8 weights) | 112 | 0 | 62 | 99.64 / 99.13 |
| `table4` (same weights, a loop over a PROGMEM table) | 86 | 2 | 127 | 99.64 / 99.13 |
| emlearn tree, depth 2, uint8 features | 232 | 10 | 177 | 92.70 / 90.67 |
Expand All @@ -87,8 +87,8 @@ One build setup, one harness, flash and RAM net of a no-model program, cycles pe
| emlearn MLP 4-4-1 (`eml_net`, float) | 4262 | 180 | 12562 | 99.64 / 99.49 |
| emlearn MLP 4-16-1 | 4550 | 564 | 33961 | 99.64 / 99.71 |

At a size near `wave4` there is no emlearn model (the smallest is 5x larger, at 90.7% mean accuracy); trees need depth 8 to reach 98.5%, at about 11x the flash and 7.6x the cycles of `wave4` for 0.7 pp
less accuracy; the MLP is the only model more accurate (+0.15 to +0.44 pp at 4-16 hidden units), at about 97x the flash and 500x the cycles for the 4-hidden-unit net, more for larger ones (soft float on a chip without an FPU).
At a size near `wave4` there is no emlearn model (the smallest is 5.5x larger, at 90.7% mean accuracy); trees need depth 8 to reach 98.5%, at about 11x the flash and 7.9x the cycles of `wave4` for 0.7 pp
less accuracy; the MLP is the only model more accurate (+0.15 to +0.44 pp at 4-16 hidden units), at about 100x the flash and 520x the cycles for the 4-hidden-unit net, more for larger ones (soft float on a chip without an FPU).
Cycles of the cells are constant; the trees vary with the path (`compare_emlearn/results.md` has min / mean / max for every model and all depths). **Read with care:** Banknote is nearly separable and says
nothing about a harder task; one dataset; emlearn 0.23.2 as documented (the `dtype='uint8_t'` option for the best-case trees, its float MLP with the 1/255 input scaling folded into layer 0, best of 5 restarts);
emlearn's fixed-point MLP is an unfinished feature upstream and its MLP `inline` method silently emits the loadable code, so no fixed-point MLP was measured (`compare_emlearn/emlearn_repro.py`).
Expand Down
2 changes: 1 addition & 1 deletion examples/static_net/check/build.sh
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ avr sugar_roll_avr_off sugar_roll_avr.cpp "1006 B / 62 B" -DSUGAR_ROLL_AT=
avr sugar_unroll_hand_avr sugar_unroll_hand_avr.cpp "1006 B / 62 B"; same "SUGAR_ROLL_AT huge == the unrolled cell by hand" sugar_roll_avr_off sugar_unroll_hand_avr
avr registers_avr registers_avr.cpp "220 B / 4 B"
avr registers_avr_flat registers_avr.cpp "220 B / 4 B" -DFLAT # the hand-indexed twin that reads before it writes: same size
avr banknote_avr_check banknote_avr_check.cpp "2346 B / 26 B" # the Banknote cell on a board: it reports over UART
avr banknote_avr_check banknote_avr_check.cpp "2344 B / 26 B" # the Banknote cell on a board: it reports over UART
if have simavr && have python3; then
echo "== AVR: the same rows through the same function, simulated ATmega328p against the host, row by row"
avr-g++ -std=c++17 -Os -mmcu=atmega328p -DSIM_ONCE "${INC[@]}" sonar_rows_avr.cpp -o "$W/rows.elf" 2>"$W/err" || bad sonar_rows_avr "$(grep -m1 error "$W/err")"
Expand Down
2 changes: 1 addition & 1 deletion examples/static_net/compare_emlearn/results.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
| model | flash net (B) | RAM net (B) | on-device ok | cycles min / mean / max (net of null) | accuracy fold 1 / 5-fold mean (%) |
|---|---|---|---|---|---|
| null | 2342 (total) | 34 (total) | 28/274 | 19 / 19.0 / 19 | - |
| wave4 | 44 | 0 | 274/274 | 25 / 25.0 / 25 | 100.00 / 99.27 |
| wave4 | 42 | 0 | 274/274 | 24 / 24.0 / 24 | 100.00 / 99.27 |
| lin4 | 112 | 0 | 273/274 | 62 / 62.0 / 62 | 99.64 / 99.13 |
| table4 | 86 | 2 | 273/274 | 127 / 127.0 / 127 | 99.64 / 99.13 |
| eml_mlp_h2 | 4214 | 132 | 270/274 | 7978 / 8711.6 / 9145 | 98.54 / 97.89 |
Expand Down
11 changes: 8 additions & 3 deletions examples/static_net/include/waveCell.h
Original file line number Diff line number Diff line change
Expand Up @@ -24,9 +24,14 @@ namespace wave {
struct WaveOf {template<typename O> struct Part:O {
using Base=O; using Base::Base;
static constexpr u8 inv(u8 x) {return n?u8(~x):x;}
static constexpr u8 shf(u8 x) {return s>=0?u8(x<<s):u8(x>>(-s));}
template<typename I> SNET_INLINE static constexpr u8 proc(const I& in)
{return u8(u8(u8(shf(inv(u8(Src::get(in))))+p)&m)+Base::proc(in));}
// avr-gcc 7.3 -Os (measured): a right shift is done in 8 bits only when it reads a variable and nothing narrows it before
// the mask; read as a call inside the expression, or cast to u8 first, it costs a register copy (.RnD/openDerivation/godbolt/).
// A left shift is the other way round: without the early u8 it widens to 16 bits (mixed_avr_size: +30 B)
template<typename I> SNET_INLINE static constexpr u8 proc(const I& in) {
const u8 x=inv(u8(Src::get(in)));
if constexpr (s>=0) return u8(u8(u8(u8(x<<s)+p)&m)+Base::proc(in));
else return u8((((x>>(-s))+p)&m)+Base::proc(in));
}
};};
template<typename Src,bool n,int s,u8 p,u8 m> using Wave=WaveOf<Src,n,s,p,m>; // Src: a Field or Elem of the state
template<size_t j,bool n,int s,u8 p,u8 m> using RefWave=WaveOf<Ref<j>,n,s,p,m>;
Expand Down
4 changes: 4 additions & 0 deletions examples/static_net/measure/silicon_results.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,10 @@ count here is a count on the chip.
sum of cycles over 274 rows). The timed code has no interrupts, no caches and no wait states on an ATmega328p, so exact agreement is what a cycle-accurate simulator should give; this
confirms it on the chip rather than assuming it.

These are the ELFs of that date. `waveCell.h`'s `WaveOf` has since been respelled (one register copy less on avr-gcc 7.3): the `compare_emlearn`
`wave4` program now runs 43 cycles per call in simavr (`wave4.min`/`max`, 24 net of the null program) and `measure/`'s `bn_wave` is unchanged at 23;
the new ELF has not been run on the chip.

| target | number | simavr | silicon | diff |
|---|---|---|---|---|
| bn_wave | wave4 | 23 | 23 | +0 |
Expand Down
Loading