MindCap
Fast enough to use. Good enough to trust.
MindCap was a private meeting-notes app built to run its AI on the laptop: recording, transcription, summaries and chat over past sessions. Its hardest lesson arrived early. A fast transcript could still be too inaccurate to help, so the project became the work of choosing small models and proving they were good enough.
Inspect the interface
Finishing quickly was only half the job.
I tested transcription on a long recording and read the result. One route took too long; a faster one produced text that still needed substantial work. Notes could finish processing without appearing in the interface. The problem spanned model quality, waiting time and whether the user could understand what the app had done.
“it's supposed to provide the "first usable" dialogue, but it falls short of this.”
Make the models compete.
I set the constraints: run locally, fit a consumer GPU, carry a licence we could use commercially. Then I had my agents build benchmark harnesses for MindCap’s own jobs instead of trusting public leaderboards. Speech models were scored for errors and speed against human-transcribed meetings and business calls, language models ran task suites, and both were tested together for GPU memory. Every candidate followed a written promotion path: benchmark, app regression, hardware check, manual acceptance.
- 01Set the constraints
- 02Benchmark the candidates
- 03Test inside the app
- 04Record the decision
How the work evolved.
- April–May 2026 / Quality
Define what “usable” actually means.
I compared the wait with the transcript I could read and pushed back when live text was quick but inaccurate. When I asked how we would choose between models, the honest answer was that we lacked a proper benchmark, so building one became the work. Names, wording and continuity joined speed as measures of quality.
- April–June 2026 / Selection
Let the data choose the model.
On a 21-minute meeting with a human reference transcript, Parakeet scored 24% content word error rate at about 26 times real time, against 33% for the local Cohere baseline. A native parakeet.cpp build at 6-bit later reached 23%, transcribing the meeting in about eight seconds on an RTX 4070. For notes and chat, Gemma 4 E4B (4-bit) became the default candidate after head-to-head runs against Ministral 3, Qwen3 and Qwen3.5.
- June 2026 / In the app
A benchmark isn’t the product.
When a “working” speech result turned out to be mocked data, and when the language benchmarks looked too simple, I said so. Prompts, context budgets and memory windows were set per model size, then judged inside the real workflows: catch-up mid-meeting, prepared notes and chat over past sessions.
- May–June 2026 / Control
Make ambient context a deliberate choice.
As the project expanded into screen context, I required that experimental capture to be off by default. The source initialises meeting context as disabled. Local inference and context capture are separate design decisions, and both deserve clear controls.
Model integration is a product problem too.
MindCap taught me to judge the complete experience: the input, the wait, the quality of the output and how someone checks it. It also taught me to distrust a passing number until I knew what it measured. On a local product, hardware and licensing constrain the choice as much as accuracy does.
The part I owned.
Product direction, model constraints and acceptance criteria, evaluating output through use, and directing AI agents to build and run the benchmark harnesses and the application. The speech and language models are third-party; I chose, configured and integrated them, and decided what met the bar.
About these images
Drawn from April–June 2026 development conversations and the MindCap repository, re-read on 1 October: benchmark documentation, model decision records and a dated parakeet.cpp report. Benchmarks ran on one workstation (RTX 4070) with small, purpose-built fixtures. The Light Academia images are synthetic demo captures.
Status and scope
MindCap is retired portfolio work, with no active service. Results come from one machine and small fixtures, so they are not general performance guarantees, and no model reached final manual acceptance. No model was fine-tuned or trained: the work was selection, configuration, prompting and integration. Earlier versions offered cloud options, so this page makes no blanket on-device claim for every version.
Technical foundations
Tauri · Rust · sherpa-onnx · Parakeet · llama.cpp · GGUF · CUDA. The stack reflects the project’s web or desktop workflow and its integration needs.
If you are adding AI to a real workflow, this is how I work: set the constraints, measure candidates on your own data, and test the result inside the product before trusting it.
Let’s discuss itThe interface, in detail.




Samplify