User Guide
Windows
Navisual watches the app you point it at, and shows you where to click — one step at a time, out loud and on screen. It never clicks or types for you. This guide covers the parts that aren’t obvious from the window.
1. Getting started
Download and run the installer from the home page. It’s code-signed, but a newly-released app still trips Microsoft SmartScreen — click More info → Run anyway if you see it.
On first launch you’ll see a short privacy notice explaining what leaves your computer. You don’t need an account: the free tier signs you in anonymously and gives you 30 requests to try it with.
Then the loop is always the same:
- Open the app you want help with and leave it visible.
- Type what you want to do in Navisual’s box — plain language, e.g. “make a pivot table from this data” or “add a bevel to this cube”.
- Follow the pointer. Navisual takes one screenshot, works out the next step, and draws a marker on the exact button, with the instruction spoken and captioned.
- Do the step yourself, then press Ctrl+~ for the next one.
Navisual only captures a screenshot at the moment it needs one — when you ask a question, or when you ask for the next step. It is not watching continuously.
2. Choosing what Navisual sees
By default Navisual looks at whichever window is in front. The header chip shows which app that currently is — Shared: Excel, and so on. Its border flashes on screen briefly when the target changes, so you can always see what’s being shared.
Click that chip to change it. You get three kinds of choice:
- A specific window. Picking one pins it (📌): Navisual keeps looking at that window even when you click elsewhere. This is the right choice whenever you’ll be switching between windows, or when the app opens dialogs that steal focus.
- Entire desktop (or Screen 1 / Screen 2 on a multi-monitor setup) — for tasks that span apps, like dragging a file from Explorer into another program. On multiple monitors, pick the individual screen rather than the whole desktop: stitching every screen into one image shrinks the detail past the point where the AI can read it.
- Unpin to go back to following the foreground window.
When one window is shared, everything outside it — including other apps overlapping it and Navisual’s own panel — is greyed out before the image is sent. The AI sees your target app and nothing else.
3. Following the guidance
Each step gives you an instruction, spoken aloud with a caption on screen, and — when Navisual can pin down the exact control — a marker on the button itself.
If no marker appears, you’ll see a quiet note saying the pointer isn’t available. That’s deliberate. Navisual would rather show you nothing than point at the wrong thing, and some surfaces genuinely can’t be read from the outside — a 3D viewport, a drawing canvas, a game. The written instruction is still correct; only the arrow is missing.
→ Next asks for the following step. Autopilot (in the action row) does the same automatically whenever the screen changes enough to look like you finished the step. It’s off every time the app starts, on purpose — it costs a request each time it fires. Autopilot suits long, obvious sequences; manual stepping suits anything where you want to think between steps.
Anything you'd have to type is put on your clipboard. A terminal command, a URL, a filename, text for a form field — when a step involves typing, the exact text is copied for you and the step shows a 📋 badge, so you paste rather than retype it. Keyboard shortcuts are the deliberate exception: those are named in the instruction, because a shortcut has to be pressed, not pasted. Navisual never runs a command or presses anything for you — you read it, then paste it and press Enter yourself.
Tasks that span several apps work without any setup: download something in the browser, edit a file, run a command in a terminal, and Navisual follows whichever window you're working in. If you'd rather it stayed fixed on one app while you move around, pin that window (§2).
Sometimes Navisual asks you a question instead of giving a step. Type the answer in the reply box and press Enter. If it can’t see enough to be sure, the honest answer — “bring the right window forward” — is what you’ll get.
The re-analyse banner appears when the screen changed a lot while the AI was thinking, meaning the advice may describe a screen you’ve already left. Take it as a hint, not an error — if the step still makes sense, carry on.
4. Hotkeys
Hotkeys are the smooth way to use Navisual, because they never move keyboard focus away from your app. Clicking Navisual’s panel does — and Windows then eats your next click on the target app, because that click is spent re-activating the window. If you’ve ever had to click a button twice, that’s why. Use the hotkey and it doesn’t happen.
| Action | Default | What it does |
|---|---|---|
| Next step | Ctrl+~ | Ask for the next step — the one you’ll use constantly |
| Mark wrong | Ctrl+E | Tell Navisual this step is wrong (see §5) |
| Voice input | Ctrl+D | Push-to-talk: speak your request instead of typing |
| Pause / cancel | unset | Stop guidance — assign your own key |
| Toggle icon mode | unset | Shrink the panel to the goldfish icon, and back |
The last two ship unassigned deliberately, so they can’t clash with a shortcut you already use. Set your own in Settings → Hotkeys, where you can also clear any binding with the ×.
5. When it points at the wrong thing
Under every instruction there’s a ✗ Wrong button. Pressing it asks what was wrong, and the four answers do genuinely different things — picking the accurate one gets you a better fix, faster:
| You say | What happens |
|---|---|
| Wrong spot (right instruction, marker on the wrong button) | Navisual re-searches your screen locally, excluding the spot you rejected. No AI request, usually instant — the advice was right, only the aim was off. If the second try also misses, it escalates to the AI. |
| Can’t find it (nothing was marked, or the thing described isn’t there) | Same local re-search first, then a fresh look at your screen. |
| Wrong instruction (this step isn’t what you need) | Goes straight back to the AI for a new plan, told what it got wrong. |
| Already did that | Moves on. If the next step is already prepared it advances instantly, without spending a request. |
You can also just type a message while the picker is open — it’s sent as free-text feedback along with your correction.
Several boxes instead of one arrow? That means Navisual found more than one equally good match — three cells that all say “Total”, say — and it isn’t going to guess. Click the one you meant. Your click is the answer; there’s no dialog to dismiss.
6. AI providers & cost
Navisual works with several AI backends. Choose in Settings → Provider.
- Managed — free. No key, no account. 30 requests per device. Note: free-of-charge AI models generally retain requests — including your screenshot — to train on. If that matters to you, use any option below instead. (Details in the privacy policy.)
- Managed — pay as you go. Top up from $5 and Navisual routes to strong paid models for you, with no API key to manage. You buy coins (a coin is $0.005, fixed); a request costs 6, 12 or 18 coins depending on which quality tier you pick — so $5 is roughly 55–165 requests. Paid providers state they don’t train on your data.
- Bring your own key. Anthropic, Google Gemini, OpenAI, DeepSeek, Qwen, or any OpenAI-compatible endpoint. Requests go straight from your machine to that provider — our servers aren’t involved at all — and they bill you directly.
- Ollama (local). Runs on your own machine. Nothing leaves your network, nothing is billed, and it works offline. Quality depends entirely on the model you run.
Which to pick? If you’re trying Navisual out, the free tier is fine. If you’re using it on real work, pay-as-you-go or your own key — both are markedly more accurate, and both keep your screen out of training data. If your screen content can’t leave your machine at all, Ollama is the only option that guarantees it by construction.
Your balance is in the header; click it to reach billing. Under the info (ⓘ) button, Usage breaks down what you’ve spent.
7. App-specific support
Navisual reads most Windows apps through the accessibility information they already publish, which is why buttons, menus and ribbons work well nearly everywhere. Some apps need more.
Browsers are one of the best cases — web pages publish rich structure, so pointing inside a web app, form or admin panel is usually exact. Chrome and Edge get a dedicated fast path.
Terminals (Windows Terminal, PowerShell, a shell inside your editor) work through the clipboard rather than through pointing: there are no buttons to aim at, so Navisual gives you the command to paste instead (§3). It never runs one for you.
Excel, Word and PowerPoint are handled specially — Navisual can resolve a cell reference, a phrase in your document, or a shape on a slide to exact pixels, rather than guessing visually. Nothing to set up.
Blender offers an optional add-on. Blender draws its whole interface itself, so from the outside it’s just an image — the add-on lets Blender tell Navisual where its tools are, and which mode and brush you’re in. It makes tool-shelf and properties-tab guidance far more reliable.
- Navisual offers to install it when it notices you’re working in Blender. Click Install, and it copies the file into Blender’s add-ons folder.
- Then restart Blender, open Edit → Preferences → Add-ons, search for Navisual, and tick its checkbox. That tick is the consent step — Navisual can’t and won’t do it for you.
- It is read-only: it answers questions about the interface and cannot click, type, run commands, or touch your file. It listens only on your own computer and sends nothing to the internet.
- Each Blender version has its own add-ons folder, so install it once per version you use. Untick it any time to turn it off.
Where Navisual struggles: apps that paint their own interface with no accessibility information and no add-on — SolidWorks’ ribbon is the clearest example — and very small text, which on-screen text recognition can miss. In those cases you’ll still get correct written instructions; you may not get the arrow.
8. Troubleshooting
| Symptom | What to do |
|---|---|
| My first click on the app does nothing | You clicked Navisual’s panel just before, so Windows spent that click re-focusing the app. Use Ctrl+~ instead of clicking → Next. |
| No pointer, ever, in one app | That app likely paints its own interface (3D viewport, canvas, game). Follow the written instruction; see §7. |
| It’s guiding the wrong app | Click the header chip and pin the window you actually mean. |
| Requests time out | Usually a slow free model under load. Try again, or switch provider in Settings. |
| “Free requests used up” | Top up with coins, add your own API key, or run Ollama locally — all in Settings → Provider. |
| The pointer is in the wrong place after unplugging a monitor | It should re-align itself within a second. If it doesn’t, restart Navisual and let us know. |
| Voice input transcribes Chinese as pinyin | Speech recognition can’t detect a language before you speak. Set it explicitly in Settings → Audio. |
| Something else | Use Send feedback in the info (ⓘ) dialog — it fills in your version and model automatically — or open an issue on GitHub. |
9. Privacy in one screen
- Screenshots are taken only when you ask a question or ask for the next step — never continuously.
- When one window is shared, everything else on screen is greyed out before sending.
- Your screenshot goes to the AI provider you chose. With your own key or Ollama, our servers never see it.
- We don’t store your screenshots. Settings, API keys and logs stay on your machine.
- The free tier’s models may train on what you send. Paid, BYOK and local options don’t.
The full details, including the Blender add-on, are in the Privacy Policy.
Navisual is AI-assisted guidance — it can be wrong. Check anything consequential before you act on it, and never rely on it alone for financial, legal, medical or safety-critical decisions.