This is the first entry in a running series documenting the build of Pyintel Lux — an open binary telemetry standard for constrained edge systems. I’m writing this as it happens, not in retrospect.
The problem I kept running into
Every time I put an ESP32 or RP2350 into a project and wanted to know what was happening inside it, I had the same experience.
printf over UART. Strings. Formatted strings. Sometimes a CSV. Sometimes just vibes.
At some point I got serious and looked at what the “industry standard” solution was. OpenTelemetry. OTLP. Structured traces, metrics, logs. The works.
Then I actually tried to run it on an ESP32.
The smallest OTel C++ SDK build I could find was around 200 KB of Flash and wanted a heap allocator. The ESP32 I was using had 4 MB of Flash total — most of it already occupied by my application firmware and the Wi-Fi stack. And the SDK expected HTTP/2 or gRPC to send data to a collector, which meant I also needed a full TCP/IP stack active the whole time.
This is fine on a Linux server. On a microcontroller doing real-time motor control, it’s a nonstarter.
What every embedded developer actually does
Here’s what actually happens in the field, if we’re being honest:
printfover UART, parsing it by eye in a terminal.- A custom binary protocol someone hacked together in an afternoon, documented only in that one engineer’s head.
- Occasionally SEGGER RTT or Percepio Tracealyzer — which are good tools but require proprietary software and a debug probe permanently attached.
- Nothing. The developer just guesses based on observable behavior.
None of these are interoperable. None of them give you a trace from MCU interrupt all the way through to a web dashboard. And none of them are something you can open-source as a standard for others to adopt.
What I want to exist
A telemetry standard with these properties:
- Zero heap allocation — safe on bare-metal RTOS systems with no memory fragmentation risk
- Compile-time string tokenization — log strings never travel on the wire, only their integer IDs do
- 12-byte minimum frame — small enough to fit inside a CAN 2.0 frame or a single UART burst
- Transport-independent — the same frame format works over UART, UDP, CAN-bus, or ESP-NOW
- Open spec — CC0 licensed, any manufacturer can implement a compatible device
- Self-hosted — no cloud dependency, runs air-gapped on a Raspberry Pi
That’s Pyintel Lux.
Why it’s a protocol, not just a library
The thing that makes OpenTelemetry powerful isn’t the SDKs. It’s the fact that there’s a wire format everyone agrees on. An OTel-compatible backend can receive data from a Go service, a Java service, and a Python service without caring which SDK generated the frames.
That’s what Lux needs to be for the embedded world. Any device — ESP32, STM32, RP2350, a custom ASIC — should be able to emit Lux frames that any Lux-compatible receiver can decode. The spec needs to be public domain so there’s zero barrier to adoption.
The SDKs are just the first implementation of that spec.
The honest state of things today
Right now Pyintel Lux has:
- A wire frame specification — 12 bytes, CRC-16/CCITT, compile-time symbol IDs
- A C API header —
lux_emit_u32(),LUX_TRACE(), zero heap - A complete C implementation of the frame assembly + CRC
- A Python host decoder that reads frames off a serial port
- An ESP32 firmware sketch that emits frames over UART
- A 5-phase research plan with two ESP32 boards
What it does not have yet:
- Any measured benchmark numbers (the targets in GOALS.md are engineering estimates, not results)
- A production SDK
- A Rust implementation
- A working
luxdingest daemon - Any external users
This is phase 1 of a multi-phase research project. The first thing I’m going to do is plug two ESP32s in, flash the firmware, and actually measure whether the targets hold. Cycle counts. Frame rates. CRC failure rates. Stack usage. Real numbers.
What comes next
Phase 1: Flash Board A with the UART emitter firmware. Run the Python decoder. Measure lux_emit_u32() cycle count using esp_timer_get_time(). Measure maximum sustainable frame rate before buffer overflow. Write up the results here.
If the numbers are good, Phase 2 starts: swap UART for UDP and measure the same things wirelessly.
If the numbers are bad, I’ll figure out why and fix the design. That’s also worth writing about.