RunanycompatibleopenLLMonyourmachine.AnAIassistantthatneversendsyourdataout.
Download any supported GGUF model that fits your hardware. Browse normally, but keep your AI prompts, uploaded documents, and page analysis contained locally — zero AI APIs, zero cloud endpoints.
› Stated exactly, because the distinction is the product: this is still a browser, so the sites you visit see your requests as they would anywhere, and non-URL text typed in the address bar goes to Google as a search. What never leaves is what the assistant works on — page content, chat and attached files — unless you point it at an AI engine on another machine yourself, which the model picker marks as leaving it.
The competition is not
Chrome. It is the sidebar.
An AI sidebar is most useful when it sees the most, and in every other browser everything it sees leaves your computer. That is the structural problem of the whole category, not a setting any of them can flip. Nightjar Browser answers from a model running beside the browser, so there is no endpoint the page could be sent to.
This matters most where it is not a preference but a requirement: client files, patient records, privileged documents, anything under a policy that forbids sending data to a third party. A smaller model that never leaves the building is a trade some people cannot make and others must.
| Browser | Sends usage data home by default | Blocks trackers by default | Sends the open page to a cloud AI | Keeps a record of your browsing |
|---|---|---|---|---|
| Chrome | Yes (Google) | No — Safe Browsing screens malicious sites, not trackers | Yes, if Gemini features are used | Yes — history, cache and cookies, synced if you sign in |
| Edge | Yes (Microsoft) | Partial — Tracking Prevention, Balanced by default | Yes, if Copilot is used | Yes — history, cache and cookies |
| Nightjar Browser | No | Yes — EasyList and EasyPrivacy, on for everyone | No — there is no cloud AI in this browser | No — private by default: history, cache and cookies stay in memory |
That No is a claim about the code, not a policy. There is no telemetry client. Two things reach the network without being asked each time, both listed in the privacy policy and both switchable in Privacy & blocking: one check for a newer version at launch, which sends the version you have and nothing else, and the ad-blocking filter lists, downloaded from easylist.to every three days. A crash report is a file you read and choose to email. Edge Tracking Prevention is real and worth naming honestly; what it does not do is block ads.
One guarantee,
and a browser around it.
The assistant is the reason to switch. Everything below it — blocking, banner suppression, tab suspension — is table stakes that keeps the browser competitive, and none of it is a reason to switch on its own: a Chrome user running uBlock Origin already has most of it.
Ask about the page you are on
The sidebar reads the current tab — article, documentation or a PDF — and answers from it, citing which part it used. Replies stream as they generate. A source small enough to fit is sent whole and in reading order rather than as reshuffled fragments.
Ask about your own files
Drop a PDF, Excel workbook, CSV or source file into the chat and it is indexed alongside the page. Ask for a total, an average or the highest row of a spreadsheet and Nightjar calculates it over every row, showing its working, instead of trusting a small model's arithmetic.
Fill the form in front of you
A Fill button appears when a page has fields it recognises, matching them against a profile stored encrypted on this machine. An assistant with reach into credential and payment fields is exactly the one you want running locally — that argument needs no privacy lecture.
Block ads and trackers
Every request a page makes is checked, inside the browser, against EasyList and EasyPrivacy with Brave's adblock-rust engine, before it is sent. Refused ad frames and images collapse, so no blank box is left behind — and sites that ask you to switch your ad blocker off are not given a reason to. Nothing is decrypted or proxied.
Cookie banners, gone
OneTrust, Didomi, Cookiebot, Quantcast and TrustArc banners collapse on DOM creation; anything unrecognised gets Reject All clicked for you. Sec-GPC: 1 and DNT: 1 go out with it, which several US state laws require a site to honour as an opt-out.
Background tabs stop costing anything
45 seconds after a tab loses focus its renderer is frozen and its memory handed back; after 30 minutes unseen, or sooner if memory runs short, it is unloaded. Fifty tabs of Wikipedia and Hacker News: 6,253 MB all live, 1,070 MB with forty-nine frozen, 633 MB with them unloaded.
Read pages aloud, offline
Through the voices Windows already ships, so nothing is sent to a speech service to be spoken back to you.
It picks the model for your machine
Nightjar Browser sizes the model and its context from each GGUF header against your RAM and VRAM, hides models that will not run, and never asks you to choose a layer count. Models you already have in Ollama or LM Studio appear in the same picker.
Two different claims, worth separating.
Pages load faster with blocking on, mostly by not stalling. The AI is fast because it is local. Only one of those is about the browser, so they are measured separately — one machine, one network, three passes per arm. Indicative of the effect, not a benchmark suite.
The AI is fast because it is local: no network round trip, and it works on a plane. Replies stream, so text appears while it is being written — first words at 0.31 s against 5.05 s for the complete answer on Qwen 3.5 2B.
A 400-token answer — about 300 to 400 words — takes roughly 6 seconds on the 2B models and 15 on 9B. Qwen 3.5 2B is not faster than the larger Gemma E2B: its recurrent layers run less efficiently on Vulkan than plain attention does.
Same models, same machine, no graphics card. The bars share the GPU panel scale, which is the point: a graphics card is strongly recommended.
CPU-only works, and on a small model it is usable — 15 tok/s is about as fast as most people read. But 9B or 12B on CPU is two minutes per answer, which does not feel like a conversation.
On Windows, Nightjar Browser embeds WebView2 — the same Chromium runtime Edge itself renders with. The unblocked arm is therefore genuinely Edge with no ad blocker, measured back to back with the blocked run on the same machine and network.
Measured on 5 August 2026 with the filtering proxy Nightjar Browser used then; the default is now Brave's adblock-rust engine, and this has not been re-run on it. Read the controls first. Wikipedia and Hacker News have no trackers, so their 2.0% and 4.8% are the measurement noise floor — anything under about 5% means nothing, CNN included. The honest headline for Yahoo is not eight times faster every time but removes the stalls: unblocked it took 9.6 s, 1.0 s and 9.9 s; blocked, 0.9 s, 1.1 s and 1.3 s.
Speed is the easy half. Nightjar Browser also grades the model you loaded, once, against a page, a spreadsheet and a PDF whose answers are known — because a parameter count does not tell you whether a model can read a value out of a table, total two of them, or admit a figure is absent. Grading is substring and numeric matching, never a model judging a model.
Bigger is not reliably better: the 1.2 GB Qwen scored full marks and the 5.3 GB one did not — Qwen 3.5 9B answered 4,096,000 for a sum that is 7,096,000, and Gemma 4 E4B 7,196,000. Arithmetic is where every model slips. The smallest is where it matters most: Llama 3.2 1B named the wrong region for highest revenue, and when asked for a figure that was not on the page it supplied one anyway. Measured 29 Sep 2026.
A tab per WebView2 means a renderer per tab, exactly as in Chrome. What differs is what happens to a tab you are not looking at: 45 seconds after it loses focus its renderer is frozen and its memory handed back, and after 30 minutes unseen it is unloaded altogether — reloaded when you come back to it.
Opening fifty tabs quickly still passes through the 6 GB figure on a machine with memory to spare: tabs freeze 45 seconds after they go to the background. When free memory runs short they are frozen at once instead, and unloaded if that is not enough — that behaviour is described here, not benchmarked on a small machine. Unloading costs a reload, and anything the page held that its address does not. This is also not a Chrome comparison — that needs Chrome driven the same way and has not been done.
The packaged build ships Vulkan
Vulkan runs on NVIDIA, AMD and Intel alike, rather than a separate 1 GB CUDA build for NVIDIA only — and it is a 20x smaller download, 32 MB against 611 MB. Measured on an RTX 3060 12 GB with tests/benchmark_backends.py:
| Workload | CUDA | Vulkan |
|---|---|---|
| Qwen 7B generation | 63.6 tok/s | 53.1 tok/s (−16%) |
| Qwen 1.5B generation | 187.9 tok/s | 195.7 tok/s (+4%) |
| Reading a 10,000-token page | 5.0 s | 5.6 s (−10%) |
Sixteen per cent on the largest model, and nothing at all on the smallest. These were taken at a different context size from the tables above, so treat the two as separate measurements rather than one series.
The browser runs on anything.
The AI panel is what needs hardware.
Because the model runs on your computer instead of someone else’s server. Also required: .NET Framework 4.7.2+ and the WebView2 runtime, both standard on Windows 10 1903 and later. An internet connection is needed for first-time setup; everything after that runs offline.
| Minimum | Recommended | Best | |
|---|---|---|---|
| OS | Windows 10 (1903+), 64-bit | Windows 11 | Windows 11 |
| RAM | 4 GB | 16 GB | 32 GB |
| GPU | none (CPU only) | 8 GB VRAM | 12 GB+ VRAM |
| Free disk | 2 GB | 10 GB | 30 GB+ |
| Models available | 2 of 7 | 7 of 7 | 7 of 7, longest context |
RAM only — no graphics card
Computed from each model’s own GGUF header, not estimated.
| Your RAM | Available | Recommends | Context |
|---|---|---|---|
| 2 GB | none | — | — |
| 4 GB | 2 of 7 | Qwen 3.5 2B (1.2 GB) | 32k |
| 8 GB | 4 of 7 | Qwen 3.5 2B (1.2 GB) | 32k |
| 16 GB | 7 of 7 | Qwen 3.5 2B (1.2 GB) | 32k |
| 32 GB | 7 of 7 | Qwen 3.5 2B (1.2 GB) | 32k |
The largest model that fits is rarely the best choice without a GPU: a 6.6 GB model on CPU generates a couple of words per second. The recommendation accounts for that, and for whether a model would be squeezed below a useful context.
With a graphics card
For Qwen 3.5 9B (5.3 GB), a typical mid-size model. Nightjar Browser sizes layers and context to your card automatically.
| Your VRAM | Qwen 9B runs | Layers | Context |
|---|---|---|---|
| 4 GB | partly on the GPU | 16 of 32 | 8k |
| 6 GB | partly on the GPU | 26 of 32 | 8k |
| 8 GB | entirely on the GPU | 32 of 32 | 16k |
| 12 GB | entirely on the GPU | 32 of 32 | 32k — every model fits |
| 16 GB+ | entirely on the GPU | 32 of 32 | 32k on the largest models too |
The 8 GB row shows the trade-off: the whole model runs on the card, and the context is held to 16k to keep it there. Keeping every layer on the GPU is worth far more than a longer context, because a layer moved to the CPU costs speed on every single word. NVIDIA uses CUDA; AMD and Intel use Vulkan.
What it cannot do, stated up front.
These are deliberate current boundaries, not bugs to be filed. The full list, with what each gap would take to lift, is maintained honestly in the repository and is worth reading before relying on anything here.
Windows first
The Windows build is the one people use. A macOS preview for Apple Silicon (macOS 26 or newer) ships with each release; it runs, but has had far less use, and is not yet signed, so installing it takes one Terminal command. There is no Linux build.
A local 7B is not GPT-5
You are trading capability for the guarantee that nothing leaves. That is the right trade for some work and the wrong one for other work.
Retrieval matches on words, not meaning
Asking about earnings may miss a page that only ever says revenue. The model is told the page may have nothing to do with the question and to say so rather than stretch it.
Small models get comparisons wrong, and invent figures
Measured, not assumed: on the built-in 16-question suite Llama 3.2 1B named the wrong region for highest revenue and supplied a figure for something the page did not contain, and Qwen 3.5 9B got a sum wrong by three million. Ask a small model which row is highest, or to total two values, and check the answer against the page. The browser runs this suite on whatever model you load and tells you where it is weak.
The model never sees a whole long document
Only the parts retrieval selected. A summary of a 200-page PDF summarises those parts.
No OCR — scanned PDFs yield nothing
Text extraction reads a PDF existing text layer. A PDF that is a picture of a page is refused. Probed and feasible offline through the OCR engine Windows 11 already ships; deferred by decision, not blocked.
No image input, even on multimodal models
Attaching an image is refused. Text-only weights cannot take images at all; for multimodal weights the model can see but the app cannot yet send.