On-device AI at RunAnywhere (YC W26). C++ core, thin bridges to Kotlin, Swift, Flutter, and React Native.
Lately that is Wally, the CLI that runs open models on your machine, and local decision models in the SDK through llama.cpp and MLX.
ToolNeuron. Offline Android AI. Chat, images, and speech stay on the device.
Ai-Systems-New. The C++ SDK behind ToolNeuron. llama.cpp for chat, QNN and MNN for images, ONNX Runtime for speech.
llama.cpp-android. CPU-only llama.cpp for Android, with ARM kernels and big.LITTLE scheduling.
ForgeAI. Desktop app to load, inspect, and merge model files. Rust and Tauri.
I also wrote Hexagon DSP kernels that skip the QNN SDK. The benchmark reported 8 TFLOPS. The matrix unit was fused off. Writeup.
C++ llama.cpp GGML Android Kotlin Swift MLX Hexagon Rust





