Conversation
PyFloat_Pack4() and PyFloat_Pack8() avoid temporary buffer in the native byte order. Use _Py_bswapXX() functions to reverse bytes.
|
Benchmark: import pyperf, sys, _testcapi
runner = pyperf.Runner()
le = int(sys.byteorder == 'little')
for d in (1.0, float('nan')):
runner.bench_func(f'pack2({d}) native', _testcapi.float_pack, 2, d, le)
runner.bench_func(f'pack4({d}) native', _testcapi.float_pack, 4, d, le)
runner.bench_func(f'pack8({d}) native', _testcapi.float_pack, 8, d, le)
be = int(not le)
runner.bench_func(f'pack2({d}) byteswap', _testcapi.float_pack, 2, d, be)
runner.bench_func(f'pack4({d}) byteswap', _testcapi.float_pack, 4, d, be)
runner.bench_func(f'pack8({d}) byteswap', _testcapi.float_pack, 8, d, be)Result:
For example, Assembly code after: The conditional jump is replaced with more efficient |
|
cc @skirpichev |
skirpichev
left a comment
There was a problem hiding this comment.
LGTM
Though, not sure if the second case in Pack2 does make sense, see comment.
|
@skirpichev: I pushed a change to also use _Py_bswap16()+memcpy() at the end of PyFloat_Pack2(). Update benchmark results (CPU isolated, after running
Benchmark hidden because not significant (4): pack2(1.0) native, pack2(1.0) byteswap, pack4(1.0) byteswap, pack2(nan) byteswap This time it's no longer "1.11x faster", but at least, it's not slower on any benchmark: it's either faster or as fast :-)
Oh sure. It's really hard to measure the speedup, the difference is really tiny and can be lost in noise. I'm using CPU isolation on Linux to reduce the noise. Details |
PyFloat_Pack4() and PyFloat_Pack8() avoid temporary buffer in the native byte order.
Use _Py_bswapXX() functions to reverse bytes.