Editorial illustration for EchoSense-AI: why another visual assistant? Because this one runs on your phone
AI analysis / Latest briefings
From TerraNet · Our app

EchoSense-AI: why another visual assistant? Because this one runs on your phone

There are already good apps that describe the world to blind and low-vision people. Most of them do the hard part in the cloud. EchoSense-AI runs a full AI model on the phone itself, so it reads your mail, describes the room and answers your questions without sending the picture anywhere, and keeps working with no signal at all.

By TerraNet Technologies4 min read
Editorial illustration for EchoSense-AI: why another visual assistant? Because this one runs on your phone
visual assistant app
app for blind and low vision
offline ai
on-device ai
echosense-ai
Listen to this article

~4 min spoken. Keeps playing while you work in another tab.

EchoSense-AI is our own app, free on the App Store and Google Play. The product page has the full feature list.

The question we kept being asked

Why build another visual assistance app? Fair question. Pointing a phone at something and hearing what it is has been possible for years, and the newest AI models have made the descriptions remarkably good.

But the descriptive, ask-anything part of most of these apps works the same way: the photo goes to a server, a large model looks at it there, and the answer comes back. That is fine for a street sign. It is less fine for a bank letter, a prescription label, a medical appointment or the inside of your home. And it stops working in a basement, on a train, abroad without a data plan, or anywhere the signal drops, which is exactly when someone who cannot see the room most needs to know what is in it.

So we built the other kind.

The model lives on your phone

In Settings, EchoSense-AI offers two versions of Google's open Gemma 4 model to download: E2B, the smaller and faster one, and E4B, the larger and more capable one. It is a one-time download of a few gigabytes, it needs no account, and each model is pinned to a verified version before it is installed.

After that, the AI that describes a scene, reads a document, identifies a banknote and answers your follow-up questions is running on the phone in your hand. The camera image and your question are processed there and nowhere else.

Private by design, not by policy

"We take your privacy seriously" is a promise. On-device processing is a fact you can test: switch on airplane mode and the app still answers.

  • With the on-device model, no images or prompts leave the phone.
  • Camera frames and results are not kept after they are processed, and a conversation lives in memory only while it is open.
  • There is no EchoSense-AI account to create, and settings stay on the device.

If you want a cloud model for a particularly hard task, you can connect your own Gemini or OpenAI key. That is optional and off until you turn it on. Before anything is sent, the app tells you what would go (the image, the prompt, and any object or depth readings) and offers to do it on the device instead. If the network is down, it falls back to the on-device model on its own.

What that looks like in a day

Reading the menu. Point the camera at a café menu and Read speaks it, dish by dish and price by price. Instant Text reads as you pan; Guided Document Scan talks you into lining up a full page, then captures it.

A sign in another language. Translate reads a Spanish pharmacy notice — opening hours, and the number to call in an emergency — and speaks it in English.

Asking about a letter. Photograph a letter in Assistant and type or say "What time is my appointment?" The answer, "Tuesday, October 20th at 10:30 a.m.", comes from the model on the phone. You can keep asking about the same letter, and the conversation remembers what you were talking about.

Moving through a crowded room. Turn on Safety in Navigate and the app pairs object detection with the iPhone's LiDAR sensor, or ARCore Depth on Android, to call out what is close: "Caution. Person, 2 meters, right. Chair, 2.5 meters, left." Ask it to describe the scene and you get the layout, not only the obstacles.

Checking the change. Hold up a banknote and Currency says the denomination aloud and on screen, offline.

Alongside those, there are barcode and QR scanning that read links aloud rather than opening them, colour and light-level detection, and a full-screen magnifier with freeze frame, high contrast, inverted colours and a flashlight.

Safety is an experimental aid. It estimates distance and is not a certified mobility device; it does not replace a cane, a guide dog or the way you already get around safely.

Built for VoiceOver and TalkBack first

Every control is a labelled native control that works with VoiceOver, TalkBack, Voice Control and Voice Access. Every result is spoken. The buttons are large and high-contrast, and the six tools — Navigate, Currency, Read, Assistant, Translate and Settings — sit one tap apart in the same bottom menu.

Where it runs

iPhone on iOS 17 or later, and Android 12 or later on a 64-bit ARM device. The on-device model wants memory: 6 GB is recommended for Gemma 4 E2B and 8 GB for E4B, plus a few gigabytes of free storage for the download. LiDAR-equipped iPhones give the most accurate distances in Safety.

EchoSense-AI is free. Download it, fetch the model once over Wi-Fi, then try it with your phone in airplane mode.

AI Tools