What the product cannot do today, why, and what each would take to lift.
This is about capability limits. For the code-audit backlog (unverified download digests, the org-name question, the UI-thread scrape) see FIXES.md. Anything listed here is a deliberate current boundary, not a bug to be filed.
Last reviewed: 2026-08-25.
pypdf reads a PDF's existing text layer. A scanned or photographed
document has no text layer, so extraction returns empty and the attachment is
refused with:
No selectable text found. This PDF is likely a scan, and there is no OCR available offline.
A PDF exported from Word, a browser, or most report tools has a text layer and works fine. A PDF that is a picture of a page does not.
Feasibility (probed 2026-07-28, works end to end): every piece needed is already on a Windows 11 machine, with no new download:
| Piece | Status |
|---|---|
pypdfium2 — renders any PDF page to a bitmap |
already a dependency |
Windows.Media.Ocr — offline OS OCR engine |
ships with Windows 11; en-US present |
A synthetic scan (text PDF flattened to an image-only PDF) round-tripped
through render → OCR and came back verbatim, including the figure 8,450.
What it would take: an engine/ocr.py that renders empty pages with
pypdfium2 and calls the OS engine, plus a fallback hook in
extract_attachment. The one open decision is how to reach WinRT from Python:
the winsdk pip package (in-process, ~0.2–0.4s/page, adds a dependency) or a
PowerShell shell-out (no new dependency, ~1s/page of process startup). A
per-page fallback rather than all-or-nothing would also catch documents that
mix real text pages with scanned inserts.
Deferred by decision, not blocked.
Attaching an image is refused. The refusal names the loaded model, because the reason differs:
-VL) — the model can see, but this app cannot
send.The second case is the app's limit, in three places:
inference/manager.py builds the llama-server command with no --mmproj,
so no vision projector is loaded.query_completion sends "content": prompt as a plain string. Images
require the OpenAI content-parts array with image_url data URIs.inference/registry.py entries download only the main GGUF; no mmproj
companion file is fetched for any model.grep for mmproj|image_url|libmtmd across the codebase returns nothing.
Do not tell users "switch to Gemma and it will work" — it will not, until
the above is wired. engine/attachments.py:image_rejection_reason() already
takes a vision_enabled flag: flip it to True once the projector is loaded
and the messages correct themselves.
(Resolved 2026-07-28 — previously there was no memory at all.)
Prior turns are now sent with each question, so follow-ups and references
resolve. What remains is a limit rather than a gap: history competes with
retrieved documents for one window, and engine/context_budget.py divides it
— roughly 70% to retrieval, the rest to history, with a floor under retrieval
so an attached document is never squeezed to nothing.
Consequences worth knowing:
settings.json already restricts to owner-only file permissions.The window is no longer fixed at 8192; it is sized per model and machine (see §11), so on capable hardware there is far more room than there was.
(Resolved 2026-07-31 — previously the whole reply appeared in one go after a silent wait of many seconds.)
query_completion takes an on_delta callback; given one it sets
"stream": true and reads the SSE body, so text appears as it is generated.
Measured on Qwen 1.5B on CPU: first text at 0.28 s against 4.83 s for the
complete reply, and 100 update frames across a 12.4 s answer.
What remains is a deliberate limit rather than a gap. Frames are coalesced on
a ~100 ms timer (MainWindow.STREAM_FLUSH_SECONDS) instead of being forwarded
per token, because delivery goes through evaluate_javascript, which is
synchronous and pumps the Tk event loop while it waits — roughly 10 ms of
blocked UI thread per message. At 47–162 tokens/second a frame per token would
spend more time painting than generating and the browser would stop responding.
Consequences worth knowing:
_run_in_background starts a bare daemon thread
and keeps no handle, so a long reply cannot be cancelled once it starts —
it can only be ignored.engine/rag.py is a hand-written TF-IDF + cosine-similarity index. It matches
on shared words, with no notion of meaning.
Mitigated but not solved: the opening chunk of every source is always included
(anchor_indices()), so an attachment is never dropped entirely just for
missing the query wording.
What it would take: a local embedding model. That means a second model
download and a second server process, or an ONNX embedder via the already
installed onnxruntime.
| Limit | Value | Where |
|---|---|---|
| Max single file | 32 MB | MAX_FILE_BYTES |
| Max extracted text | 400,000 chars, then truncated | MAX_TEXT_CHARS |
| Retrieved context | ~70% of the free context budget | engine/context_budget.py |
The model never sees a whole document — only the retrieved chunks. "Summarize this 200-page PDF" summarizes the parts retrieval selected, not the PDF. For broad questions over long documents the answer will be partial, and it will not say so.
The budget is now measured in tokens against the model's real window rather than a fixed chunk count. The previous fixed count could overflow outright: 18 chunks is roughly 7,500 tokens against a 7,168-token budget at 8k context.
Unsupported formats: DOCX, PPTX are refused. They are ZIP containers of
XML, not text, and need a parser (python-docx) that is not a dependency.
Legacy .doc/.xls/.ppt are worse and effectively out of scope; .xls
says to save the file as .xlsx or CSV.
Supported: 49 extensions — PDF (with a text layer), HTML, TXT, MD, CSV/TSV,
Excel (.xlsx, .xlsm, every sheet), JSON, XML, YAML, and 27 source-code
extensions.
Calculations on spreadsheets. For a data question about an attached CSV, TSV or Excel file (a total, average, count, highest or lowest, difference or lookup), the model only writes a query and Nightjar computes the answer over every row, showing its working. Grouped breakdowns ("total per region") are not supported yet. When a question cannot be calculated, the answer is the model reading the table as text, and a note before it says so and why. Nightjar refuses to calculate from a file cut at the size limit, or from a workbook whose formulas were never calculated by Excel, since either would give a wrong total.
A file stays attached for the whole session, across every subsequent question, until removed from the tray, cleared, or dropped by "New Chat". This matches how chat attachments behave elsewhere, but it means a file attached for one question keeps consuming the retrieval budget for later, unrelated ones.
Linux: none, and no near-term path to one. The browser is WebView2 via
.NET interop (pythonnet, tkwebview2) on Windows and WKWebView on macOS.
Neither has a Linux equivalent here, and nothing is planned.
macOS: Apple Silicon only, and much less exercised. There is a build --
build_mac_app.py, produced by .github/workflows/build_macos.yml -- with
the inference runtime sealed into the bundle and Metal for acceleration.
Intel Macs are excluded deliberately: no unified memory and no Apple Silicon
GPU means CPU inference at single-digit tokens per second.
Do not read that as parity. The Windows build is the one with users on it; the macOS build has had far less real use, and features are likelier to be missing or half-wired there. A separate PySide6/QtWebEngine port that would unify both platforms is planned but not started.
The OCR path in §1 is absent everywhere, and would be Windows-specific if it existed.
The governing rule: always show the video, even if that means showing the ad. Blocking may fail by letting an advert through. It may not fail by costing the user the content they came for.
Which engine this describes. The default blocker is Brave's adblock-rust,
matching every request against EasyList and EasyPrivacy inside the browser; it
applies this rule through video_safety_exceptions(), layered over the lists
so the exceptions win. The steps below are the hostname blocker's, which is
the fallback when adblock-rust cannot load -- the same rule, spelled out.
is_blocked() enforces it in three steps:
allowlist.txt / DEFAULT_ALLOWLIST entry wins outright.api., cdn., manifest., registry., hls.,
or a CDN token in the domain such as phncdn), is blocked only if our own
curated list names it. A downloaded community list is not trusted here.Step 2 exists because registry.api.cnn.io ships in the StevenBlack hosts list
as a tracker and is also the API CNN's player calls to resolve a video to its
media. Blocking it did not remove an advert; it meant CNN videos never played.
Note it is on cnn.io while the page is cnn.com — a sibling domain, so
first-party matching alone does not save it, which is why the infrastructure
heuristic is there too.
The cost, stated plainly: a tracker hiding behind an api. or cdn. label
that only a community list knows about will load. This is mitigated by keeping
known analytics and identity vendors in the curated list (amplitude, mixpanel,
segment, hotjar, clarity, permutive, lotame and ~30 more), since curated rules
still apply to infrastructure-labelled hosts. Verified: 27 ad/tracker hosts
blocked, 0 leaks; 15 content-delivery hosts allowed, 0 breakage.
Four further protections come from the same rule:
block_ssai_hosts (Settings → Privacy & blocking → "Block
server-side ad-insertion hosts") turns them on for stronger pre-roll
blocking, at that risk. The setting is authoritative:
when off, these hosts are allowlisted, because the downloaded list contains
them too.v.redd.it, TikTok CDNs,
Dailymotion edges, gcore, bitmovin, mux, kaltura, wistia, b-cdn.net), since
no heuristic recognises them and the weekly blocklist refresh could otherwise
start blocking them silently.-ad- cosmetic wildcards are guarded with :not(:has(video)), so a
wrapper whose class merely contains -ad- is never hidden if it holds the
player. Verified :has() is supported in the shipped WebView2.<video> element
serves both advert and programme there, so if the ad classes are ever wrong,
the length cap stops it fast-forwarding real content.What this still cannot fix: a CDN hostname that reveals nothing and is not on
the explicit list (ev-h.phncdn.com was caught by the phncdn token, but
something like x7f2.example.net would not be). If a site breaks, the
blocked-domain list in the shield drawer names the culprit and allowlist.txt is
the remedy.
Related failure mode, now mitigated: the filtering proxy carries all
browsing, not just adverts, so a dead listener meant every page failed with a
proxy error. The accept loop used to break on any OSError. It now retries and
rebinds the same port (which cannot change — it is fixed in WebView2's launch
arguments), the shell re-checks it on a timer, and a matcher exception fails open
rather than dropping the request. A bug in blocking should cost an advert, not
the web.
Text-to-speech works offline with no download, through System.Speech and the
SAPI5 voices Windows ships. It reads AI replies, the page selection, or a whole
article.
The limit is voice quality, and it is worse than it first appears. Windows
keeps voices in two separate registries, and System.Speech only reads one of
them:
| Registry | Voices (measured on a Win 11 machine) | Reachable today |
|---|---|---|
Speech\Voices — SAPI5 |
David Desktop, Zira Desktop | yes |
Speech_OneCore\Voices |
David, Mark, Zira | no |
The "Desktop" voices are the oldest and most robotic ones Windows ships. The OneCore set — what Narrator uses — is noticeably better, already installed, and invisible to us.
Installing Windows' natural (neural) voices does not help either. They install into the OneCore registry too, so Settings → Accessibility → Narrator → Add natural voices would leave Nightjar Browser's voice list unchanged. An earlier version of this document claimed they would "appear through the same API with no code change"; that was wrong, and only checking the registries showed it.
Reaching any of them means moving to WinRT Windows.Media.SpeechSynthesis,
which enumerates OneCore. The cost is that WinRT returns an audio stream
rather than playing it, so playback control — today free from System.Speech —
has to be rebuilt around a media player. Pause mid-sentence is the part that
gets hard; stop-between-chunks stays easy because the text is already chunked.
Not done because the current voices are adequate for the use case. Piper remains the real quality jump; neither Windows engine reaches it.
Also worth knowing:
Piper (ONNX) remains the upgrade path if the built-in voices prove too grating:
onnxruntime is already a dependency, and voices are 25–60 MB each. It was not
taken because the Windows engine needs no download at all.
The context window is sized per model and machine rather than fixed, so the same build behaves very differently across hardware. Measured from the real GGUF headers:
| model | 6 GB VRAM | 8 GB VRAM | 12 GB VRAM | 16 GB RAM, no GPU |
|---|---|---|---|---|
| llama-3.2-1b | 32k | 32k | 32k | 32k |
| qwen3.5-2b | 32k | 32k | 32k | 32k |
| qwen3.5-4b | 32k | 32k | 32k | 32k |
| gemma-4-e2b-it | 32k | 32k | 32k | 32k |
| gemma-4-e4b-it | 8k (86% GPU) | 24k | 32k | 32k |
| qwen3.5-9b | 8k (81% GPU) | 16k | 32k | 32k |
| gemma-4-12b-it | 8k (62% GPU) | 8k (88% GPU) | 32k | 32k |
(16 GB of system RAM throughout; 32k is the automatic ceiling, and a larger window can be chosen in settings where the memory allows.)
What drives the spread is mostly the weights, not the cache: every model in the catalogue keeps a whole-context cache in only a few of its layers. Qwen 3.5 makes three layers in four recurrent, and Gemma 4 limits five in six to a window of 512–1024 tokens, so Gemma 4 12B's cache is 16 KB a token where a plain design of the same size would need several hundred. The planner reads this from each file's header and charges the windowed layers as a fixed cost; the figures were checked against llama.cpp's own allocation.
Consequences: on a 6 GB card the mid-size models run partly on the CPU and
slow down accordingly, and --cache-type-k q8_0 (the "longer context on small
GPUs" lever) trades a little quality for roughly double the window.
The test suite (783 passing) is unit and integration level. It exercises Python
logic and parses ui/sidebar.html as text; it does not drive the real WebView2
control.
Consequences seen in practice: a sidebar reply handler read an undeclared
variable and threw on every model reply, rendering the chat silently dead,
while the whole suite stayed green. tests/test_sidebar_chat_render.py now
guards that specific class of bug, but the general gap remains — a change to
sidebar JavaScript is not covered until someone runs the app.
The assistant reads whatever page is open, so whoever controls a page controls
part of the prompt. engine/injection.py and the probe in shell/tab.py take
text a person cannot see out of the model's context and report what was
removed: CSS-hidden passages the browser confirms are invisible, and the
invisible-character families (zero-width, bidi overrides, the Unicode Tag
block). Instruction-shaped phrasing in visible text is reported and
deliberately left alone.
This raises the cost of the attack. It does not stop it, and nothing in any description of this product should say that it does. What is known to get through:
querySelectorAll does not reach;alt, title and aria-label, since the walk reads text nodes only;for_agent DOM serialisation, which /crawl and form detection read and
which is not swept at all.The thresholds are guesses that no attacker has yet pushed on: passages under 12 characters ignored, 40 findings per page, font sizes between 4px and 8px allowed through, and a colour-distance floor of 24.
TODO.md item 11 tracks the work. The honest description is "hidden text is removed where it can be found" -- never "protected against prompt injection".
Being told is now the exception rather than the rule, changed on 2026-08-28 after the notice was seen firing on ordinary sites. On a web page nothing is said in the conversation at all: the removal happens and a line goes to Activity. Only an attached file speaks up, and only for findings that read as an attempt to steer the answer.
The reasoning, since it reverses an earlier decision. The probe recognises
techniques, and the techniques are shared -- clip: rect(0,0,0,0) and
left:-9999px are how an accessible site writes "Skip to content" and how an
attacker hides a sentence. nytimes.com produced forty-one findings and no
attack. A warning above every reply teaches people to dismiss warnings, and
the one that matters would have arrived looking like the forty-one before it.
The cost is stated plainly: a page that attacks the assistant now does so
without the user being told in the conversation. What protects them is the
removal, not the notice. Where the removal fails -- a passage found but not
matched in the extracted text, removed: False -- there is now no visible
sign of it, and that is the gap this trade opens.
Blocking an advert stops it loading; the box the page kept for it is a separate problem. Nightjar Browser collapses frames and images whose request was refused and applies each site's own hiding rules, which clears most of it.
It does not, by default, apply the filter lists' general hiding rules --
the ~9,000 like ##.ad-slot that hide anything with an ad-like class name on
any site. Sites that check for ad blockers plant bait elements carrying those
names, see them hidden, and show a "disable your ad blocker" wall: measured on
foxnews.com, the wall appeared with them and not without them. Brave's
standard mode makes the same choice.
They are available as Hide more ad space (aggressive) in Privacy & blocking, with that cost stated beside it. On macOS neither the collapsing nor the general rules are wired up yet (see MACOS_HANDOFF.md).