the docs

How Arock works

Arock is a rock that sits on your Mac's desktop. You talk to it, or text it, and it works your apps for you. Under the rock there is no single giant model doing everything. There are seven parts, each picked for one job, and most of them are small and quick.

The earSilero VAD + Parakeeton your MacJevpicks the next step~0.2 sThe handsmacOS accessibility~1 msThe voicePocket TTSon your MacMercuryfills in the details~0.5 sGroqlooks at the screenwhen neededTabPFNthe odds, from your pastwhen stuckwordsdo thisit happenedthe detailscan't read it? lookthe replystuck?

A request goes like this. The ear turns your voice into words on your Mac. Jev reads where things stand and picks the next step. If the step needs details (which file, what to type), Mercury writes them. The hands do it through macOS accessibility and check that it really happened. Jev looks again and picks the next step, until it decides the request is done. Then the reply is written and the voice says it, also on your Mac.

Two parts only show up when needed. Groq looks at the screen when a window can't be read the normal way. TabPFN gives odds from your own history when a step went wrong or there are too many buttons to choose from.

51 of 60desk tasks passed
0safety failures
1.9 smedian short task
$0.0006median task cost

From the desk eval: 60 fixed tasks across apps, browsers, files, settings, long documents and two-app jobs, graded by a separate checker that looks at the Mac afterwards. September 2026, development build.

The ear and the voice

Everything about sound happens on your Mac. Nothing you say is streamed anywhere to be heard.

  • Silero VAD notices when you are talking and when you have stopped.
  • Parakeet TDT 0.6B (NVIDIA, int8) turns the whole turn into words in 0.2 to 0.4 s on the Mac's CPU. It replaced a smaller model after a test on 160 real recordings: on meeting audio from one distant microphone it got 39% of words wrong where the other got 71%.
  • Pocket TTS (Kyutai) speaks the reply, a sentence at a time, so the first sentence plays while the rest is still being made.
  • All three run through one pinned build of sherpa-onnx.

Knowing when you're done talking

A short pause inside a sentence is a breath, not the end. Arock waits for 700 ms of quiet before it treats your turn as over. It doesn't spend that wait doing nothing: as soon as you go quiet it reads what you said and starts deciding. If you carry on talking, that early start is thrown away as if it never happened, and nothing has been done on the Mac. If you don't, the answer is already on its way.

If you talk over the rock, it stops, and it remembers only the part of its answer you actually heard.

Jev, the decider

Jev (typesafe/jev-1.13, through OpenRouter's decisions API) is the part that is asked the most, so it is the part that has to be quick and cheap. It does not chat. It answers typed questions: here is the state of things, here are the options, which one? It returns a choice with a probability for every option.

191 msmedian decision
$0.00006per step

At each step the options are Arock's tools (open an app, press a control, set a value, use a menu, read a file, search the web, look at the screen…) plus answer, ask and think. Getting the probabilities, not just the winner, matters twice:

  • Close calls get a second opinion. When Jev's top two choices are within 10 points of each other, Mercury is asked to pick between just those two, thinking hard. In one eval run this fixed a request where "answer" beat "hide the window" 38% to 37%.
  • Nothing caps a request. There is no step limit. If Jev repeats itself, it is shown that the same call gave the same result; if it is stuck, think is one of its options.

Mercury, the writer

Mercury (Inception) is a diffusion language model: it drafts a whole passage at once and refines it, instead of writing one word after another. For short structured text that makes it fast, about half a second.

Arock uses it for the parts that need words but not opinions:

  • Filling a step. Jev says "open_app"; Mercury writes {"app": "Google Chrome"}. Lua checks the result against the tool's arguments before anything runs.
  • Writing a question when Jev decides to ask you something, as a small form with choices.
  • Breaking ties for Jev, as above.
  • Keeping notes: what to remember about you, and a one-line summary of each conversation.

The spoken reply itself comes from DeepSeek V4.1 Flash with thinking off, on whichever provider answers fastest (about 0.47 s). If it fails, Mercury writes the reply instead.

Groq, the eyes

Most of the time Arock never looks at a screenshot. It reads windows through macOS accessibility, which gives every button and field with its name, the way a screen reader does. That is faster and more exact than looking at pixels.

Some apps draw everything themselves and give accessibility nothing to read. Then Arock takes a picture of that one window and asks a vision model what's in it. That model runs on Groq (Llama 4 Scout), chosen for speed. When Groq isn't set up, Cerebras or OpenRouter runs a Qwen vision model instead. On a labelled test, the Cerebras model got 47 of 48 claims about eval screenshots right, at 0.76 s each.

A frame is that window's own pixels, not the whole screen, unless the whole screen is what was asked about.

TabPFN, the hunch

Every step Arock takes is written down with how it turned out: worked, broken, or had no effect. That makes a table. TabPFN (Prior Labs, v3.5) is a model built for tables: give it rows with known outcomes and new rows without, and it gives odds for the new ones in a single pass. It learns from the table you hand it at that moment. Nothing is retrained.

requestappstepoutcomerename a fileFinderset_valuebrokenrename a fileFinderpressworkedreply to AnaMailputworkedrename a fileFinderkeysno effectpress74%keys17%set_value9%readodds for the next step

It is asked at two checkpoints, and only there:

  • After a step broke or did nothing. It ranks the tools for the next step from what has worked before in similar spots, and Jev sees the ranking.
  • When a window has more than 10 controls and Jev is about to press one. It narrows the list to the likely ones.

A prediction can take a few seconds, which is why it is kept for when things are going wrong, not for every step. Its predictions are scored as outcomes come in, so how often it is right is measured, not assumed. It runs on a daily budget under the free tier.

The hands

The hands are Swift code in the app calling macOS accessibility: press a control by name, set a field's value, choose a menu item, send keys to one app. A move takes about a millisecond. Keys and clicks go to the app's own process, so the window you are typing in keeps your keyboard.

After a change, a witness checks that it actually happened: the app opened, the window changed, the file moved. If it didn't, the step is recorded as having no effect, and Jev is told so instead of being told it worked.

What it does on its own, and what it doesn't

Just does itreading changes nothing· read a window· list files· read a web page· check the volumeAsks firstanything that changes· move a file to the Trash· send, post or buy· change a setting· type into a formHands it backyours, always· passwords· two-factor codes· CAPTCHAs· signing in

Reading is free. Anything that changes something waits for your yes, and a no is final for that request. Passwords, two-factor codes, CAPTCHAs and sign-ins are refused by the hands themselves, not just by a rule in a prompt: a secure text field can't be typed into, and the rock hands it back to you. It never reads saved passwords.

Rocks, memory and privacy

You can have more than one rock, each with its own name, look and memory. Each rock keeps two SQLite files of its own on your Mac: what it has learned about how your Mac behaves, and what it remembers about you. One rock never reads another's files, and a test checks that nothing leaks between them.

One rock is pinned to the desktop and answers out loud. The others live in the Rocks window, where you text them like friends. Texting and talking to the pinned rock are the same conversation.

Every call Arock makes to a model is written to a trace on your Mac with its time and cost, kept for 30 days. That trace is where every number on this page comes from.

Lua inside

Almost all of Arock is portable Lua: the agent, the turn-taking, the settings, the rock's shape. A thin Swift and Metal wrapper draws the rock and holds the parts that need the operating system (audio, accessibility, capture). The rock itself is generated by a Lua script from a seed, and the same script runs in your browser when you make a rock on this site, so the rock you make here is the same one the app draws.

The same Lua runs unchanged on LuaJIT in the app and inside the Erlang VM on Arock's server, where rocks will be able to keep working while your Mac sleeps. There, a sleeping rock is just its SQLite file in storage. In testing, one wakes from storage in under half a second, and a hundred awake rocks use about 25 MB. The server is not public yet.

Why it's fast, and why it's cheap

Most of a computer task is small decisions: which app, which button, done yet? Big assistants send each of those to one enormous model and wait for it to think. Arock sends them to Jev, which answers in a fifth of a second for about six thousandths of a cent, and calls the bigger models only for the parts that need words.

0 s0.5 s1 s1.5 s2 s2.5 sA questionearJevreplyvoicetalkingOpen an appearJevMercuryhandsJevreplyvoicetalkingyou stop talking
  • Sound stays local. No upload before anything can start, and no per-minute speech bill.
  • No screenshots by default. Accessibility text is small, exact and instant.
  • The slow parts are rare. Vision and TabPFN only run when the quick path fails.
  • It starts before you finish. Deciding begins as you go quiet, not after the pause.
$0.0006median task
$0.0019most expensive 5%
~1 msa hand move

Median and 95th-percentile cost per task from the September 2026 desk eval, counting every model call a task made.