Before you add on-device AI, pick the target device
On-device AI can reduce API calls, but it does not remove model files, memory limits, or platform support. Here is what to check before adding it to a product.

When adding AI to an app, it is tempting to start by picking a model. First answer a different question: where does the app need to run?
The runtime and models available to you change between a browser, a phone, and a desktop. NobodyWho, introduced on GeekNews, is an inference engine for running LLMs locally in apps and games. Its README lists bindings for Flutter, Python, Godot, Kotlin, Swift, and React Native.
The phrase ‘on any device’ is not enough to set a release plan. The README says there is no web export and Windows ARM64 is not supported yet. The Godot binding does not export to iOS, so an iOS app needs a different binding.
Memory is more than the model file size. The project gives a rough desktop guide of about 1.5 times the model file in free RAM, or twice that on a busy machine. For mobile it suggests around twice the model file size. These are maintainer rules of thumb, not performance results for a particular app.
Check the network path for the first model load too. The README says models can be downloaded from Hugging Face or a URL and cached on first use. Inference may run on-device, but if the model is not already installed, the product still needs a download path and a clear failure state.
A planning sheet needs more than ‘runs locally’: supported operating systems and frameworks, model size and available memory, first-download behavior, update timing, and what happens when the model cannot load.
The README cannot tell us the response time or battery use in a real app. It also does not establish how much faster a particular GPU will be. Those numbers need to be measured with the target device and model.
So the first on-device AI question is less ‘Which model?’ and more ‘Which devices must we support?’ Once that is clear, model size and runtime become practical choices.

