[FACE] person_a
[OCR] CAFÉ NOIR
Real-time scene tagging through the camera. Object, face, OCR detection at 24 FPS on Apple silicon.
A media AI app that round-trips every frame to a cloud GPU is slow, expensive, and bad for privacy. We push as much inference as possible to the device: Apple’s Neural Engine, Qualcomm Hexagon, the NPU on Tensor, Dimensity, and Snapdragon X.
Real-time camera filters, on-device transcription, on-device reframing, on-device search. The cloud is reserved for heavy lifts: long renders, large-context reasoning, multi-tenant catalog work.
The result is apps that work in airplane mode, behave on cellular, and don’t burn a GPU bill on every tap. The cost curve looks more like a hosting bill than a metered API one, which is what makes consumer pricing land.
Either as standalone product apps for content owners, or as SDKs embedded inside an existing OTT app. We can ship in your design system or bring our own.
Real-time scene tagging through the camera. Object, face, OCR detection at 24 FPS on Apple silicon.
A pocket version of Kladon Clips. Drop a long-form video, get publish-ready 9:16. Editor sign-off in the elevator.
Live transcription tuned for Korean, Japanese, and Mandarin dialogue, with English bridging. On-device Whisper distill plus regional ASR fallbacks.
Swift, SwiftUI, AVFoundation, CoreML, Vision, ARKit, MLX. Apple Silicon-first performance. TestFlight from week 4.
Kotlin, Jetpack Compose, CameraX, NNAPI, TFLite, MediaPipe, ONNX Runtime Mobile. Play Console from week 4.
Where it makes sense (Flutter, React Native, Kotlin Multiplatform), we bridge to a shared inference core. Where it doesn’t, we go full native. We don’t ship ports.
Drop a Kladon SDK into your OTT or social app: scene tagging, transcription, highlight detection. You ship the feature without owning the inference plumbing or the GPU bill behind it.
On-device AI removes the cloud round-trip and the per-call inference bill. With the model sitting where the camera is, the experience is faster, more private, and dramatically cheaper to operate at consumer-app scale.
Live tagging, segmentation, face and OCR detection through the viewfinder at 24+ FPS. No buffering, no upload.
On-device speech-to-text in the user’s language, even in airplane mode. Korean, Japanese, Mandarin, and English ship out of the box.
Frames never leave the device unless the user asks. Useful for medical, enterprise, journalism, and consumer-privacy positioning.
A flat-cost mobile experience instead of a metered cloud one. Margin per active user looks like a hosting line, not an API one.
Native binaries.
On-device weights.
Shipped.