Skip to content

Requirements and speed

What jhint needs from your computer, and how fast answers actually arrive.

What it needs

  • Permission to hear the call — see Permissions.
  • Disk space for the models you choose: about 2.5 GB for the recommended built-in model, plus a few hundred megabytes for speech recognition.
  • Memory while the built-in model runs: roughly 4 GB for the 4B model, 7.5 GB for the 8B one. Nothing is reserved when no session is running.

macOS 14 or later, on an Apple Silicon Mac. An Intel Mac won’t run it, and neither will a virtual machine without the Neural Engine. The permission to hear the call is called Screen Recording.

Windows 11 (22H2 or later), 64-bit. There is no separate ARM build; on an ARM machine it runs through emulation. We have not tested jhint on Windows 10 and do not claim it works there.

The built-in model needs a discrete graphics card. This is the one requirement worth reading before you buy, because it decides whether that model is offered to you at all. On the machine we measured it on, a laptop RTX 4060 wrote about 79 tokens a second; the same machine’s processor managed 14, and its integrated graphics 11.9. Below about 15 the answer arrives slower than the conversation moves, so jhint does not offer the built-in model on a machine without a discrete card — rather than offer it and let you find out during a call.

Without a discrete card jhint still works. What changes is where the answer comes from: an external provider (Groq, Anthropic, OpenAI, Google, OpenRouter, any OpenAI-compatible endpoint) or a model you run yourself in Ollama or LM Studio. See How jhint answers.

Speech recognition works on any machine, with or without a card — it just uses a smaller model on the ones without. jhint measures your computer on first run and chooses for you.

Memory. The built-in model needs its own memory on top of what the system is using: 4 GB for the 4B model, 5.5 GB for the one that reads screenshots, 7.5 GB for the 8B model, and jhint leaves 2 GB for the system on top of that. In practice: 8 GB of RAM is the floor for the 4B model, 16 GB for the 8B one.

Disk space. The app is about 305 MB. Speech recognition adds 147 MB or 465 MB depending on the model chosen for your machine, and the answer model 2.5 GB (4B), 3.3 GB (screenshots) or 5.0 GB (8B). A typical set is around 3 GB.

Installation goes into your own user folder and does not ask for administrator rights. jhint listens to the audio your computer plays; the microphone is not used, and Windows is not asked for any permission at all.

Speed

On an M1 Max with the recommended built-in model, the first word of an answer arrives in about half a second, and the answer is written out at roughly 55 words’ worth of tokens a second. In practice: you press Answer, glance at the camera, and the answer is there when you look back.

On a laptop with a discrete RTX 4060, the recommended built-in model writes at about 79 tokens a second — faster than you can read it. That number is what a discrete card buys you: the same laptop’s processor managed 14, which is below the point where the answer keeps up with a conversation.

Speech recognition does not need a card at all — it only changes model. With a discrete card jhint uses the larger speech model; without one it uses the smaller one, which turns a window of speech into text in about 0.7 seconds. The larger model on a processor alone would take about 2.3 seconds, which is why it isn’t offered there.

The first session after launch adds a few seconds while the engine starts. It stays warm for five minutes after a session ends, so back-to-back calls don’t pay that cost twice.

An external provider replaces local speed with network speed — usually between half a second and a couple of seconds, depending on the provider and the model.

Choosing without a licence

The download page and the trial let you measure this on your own hardware rather than trust our numbers: the whole app works during the trial, on your computer, with your calls.