Browser and Local ML
Running ML models directly in the browser or on the local machine without a server enables privacy-preserving, low-latency inference.
Definition
TensorFlow.js converts models to a web format and runs inference in the browser using WebGL or WebAssembly (WASM) for acceleration.
ONNX Runtime can run models locally on CPU or GPU, enabling serverless and edge deployment.
Intuition
Browser-based ML keeps data on the user's device — no server round-trip, no privacy risk for sensitive inputs.
Smaller models (quantized, distilled) are needed for real-time browser performance — there's a size/speed/accuracy tradeoff.
Worked example
A sentiment analysis model running entirely in the browser analyzing text as you type, never sending data to a server.
A pose estimation model in the browser using the webcam for real-time movement tracking.
The math
Quantization reduces model size by using int8 or float16 weights instead of float32: typical for 4x size reduction with minimal accuracy loss.
Knowledge distillation trains a smaller "student" model to mimic a larger "teacher" model's outputs.
In practice
Privacy-sensitive applications (health, finance) where data cannot leave the user's device.
Offline-capable applications, reduced server costs, and reduced latency for real-time applications.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.