OtōtoOtōto

Your coding agent's
little brother

Otōto is an MCP server for Claude Code and OpenCode. Your agent hands it a question; a small model on a server you choose explores the code and brings back a short answer, with the lines that back it checked and quoted.

curl -fsSL https://ototo.sh | sh

macOS on Apple silicon and Linux · See the script · How it works

Claude Code asksotōto

Where does a failed request move on to the next model server?

Otōto answers

Llm::request tries each endpoint in turn; one that fails is marked down for 60 s and the next is asked. When every failure was one that passes (a dropped connection, a 502), it goes round once more after 2 s.

  • src/ototo/agent/llm.rs:144for round in 0..2 {
  • src/ototo/agent/llm.rs:163*b.down_until.lock().unwrap() = Some(Instant::now() + DOWN_FOR);
  • src/ototo/agent/llm.rs:174tokio::time::sleep(RETRY_AFTER).await;
4 turns · 6 tool calls · 20.0k tokens on the small model · 18 s
Fig. 1Real asks on public repositories and on Otōto itself, and a plugin; replies shortened. Claude Code reads a paragraph and a few quoted lines, not the files they came from.

01How it works

Your agent asks. Its little brother does the reading.

Every file an agent opens, every search result and every dead end stays in its context for the rest of the session, and each later turn carries it again. Otōto does that exploring in a context of its own, thrown away after each call.

1

Claude Code asks

A whole question to ask, or locate and callers for where something is and who uses it.

2

Otōto explores

A small model on your server reads, outlines and searches the repository with read-only tools, as many turns as it takes.

3

A checked answer comes back

A few lines, with every path:line checked against the files and quoted. A claim the code does not back is flagged.

Your agent's context, when it explores itselfEverything it read stays, and goes with every later turn.

about 2,600 lines of files and searches

Your agent's context, when it asks OtōtoThe answer stays; Otōto's reading was thrown away.

a paragraph and three lines
Fig. 2What one lookup leaves in your agent's context: an illustration, not to scale. With a paid model the smaller context is a smaller bill; with one model on both sides (OpenCode on a local model), a main agent that stays on the task.

02Tools

Some questions need a model. Most lookups don't.

Three tools hand the work to the small model. The rest answer at once, from the files and from git, for when your agent already knows what it wants to see.

ToolWhat it doesFor example
askmodelA whole question, answered in a few lines with checked citations.how is auth wired into the router?
locatemodelWhere something is defined, or where a behaviour is decided.where are retries configured?
callersmodelWho calls or uses a function, type or setting.who uses DOWN_FOR?
readCode by address, many at once: a declaration by name, a range, a config key.src/ototo/agent/llm.rs#brief, pom.xml#spring-core
outlineA file's or a directory's declarations, with their lines.src/agent/
searchExact text or a regex across the repository, as path:line hits.-F RETRY_AFTER
changesWhat a branch, a commit or a range changed, by declaration.--commit 907d144
historyThe commits behind some lines or a declaration.src/ototo/agent/llm.rs#request
digestA build or test run cut to its failures, errors and summary.ototo digest -- cargo test
$ ototo read src/ototo/agent/llm.rs#brief
── src/ototo/agent/llm.rs:290-301  brief · fn brief(e: &anyhow::Error) -> String
290| fn brief(e: &anyhow::Error) -> String {
291|     if let Some(r) = e.downcast_ref::<reqwest::Error>() {
292|         if r.is_connect() {
293|             return "unreachable".into();
294|         }
295|         if r.is_timeout() {
296|             return "timed out".into();
297|         }
298|     }
299|     let s = e.to_string();
300|     if s.chars().count() > 120 { … } else { s }
301| }
Fig. 3One function by name, with its lines: no file to open, no model involved.
$ ototo changes --commit 907d144
── changes in 907d144 changes: one commit, or a range, told by declaration, 3 files, +123 −20
M src/ototo/main.rs  +13 −2
  changed  Command (69-291)
  changed  main (361-585)
M src/ototo/mcp.rs  +19 −7
  changed  instructions (31-86)
  changed  tools (220-380)
M src/ototo/tools/git.rs  +91 −11
  added    commit (171-181)  added    range (184-205)
Fig. 4A commit told by the functions it touched, instead of a diff to read through.

03Plugins

For the files your build tools understand, and your agent has to guess at.

A plugin reads one kind of file the way its tool does: read pom.xml#spring-core returns the version Maven resolves, not the line that says ${spring.version}. Plugins are WebAssembly, signed and sandboxed: a file plugin sees the repository, read-only; a forge plugin reaches only the hosts you grant it and never sees its token.

maven file

A dependency's effective version and scope, through parents, BOMs and properties.

pom.xml
spring-config file

A property as each profile sees it, and where each value is set.

application*.yml, .properties
terraform file

A resource with its variables worked out, per environment.

*.tf, *.tfvars, *.tofu
kube file

A Kustomize resource as it builds; a Helm value and the templates that use it.

kustomization.yaml, values.yaml
gitlab-ci file

A job as GitLab runs it: includes, extends and anchors resolved.

.gitlab-ci.yml
compose file

A service as docker compose runs it, overrides and .env applied.

compose.yaml
gitlab forge

The branch's merge request, a pipeline by stage, a failed job's log cut to what failed.

gitlab.com, or yours

04Dashboard

See what it does, call by call.

ototo ui is a dashboard over Otōto's run log, on your machine: every call from every session, how long the model server took, what the small model did turn by turn, and the exact reply your agent got. ototo tail follows calls as they run. The page listens on 127.0.0.1 only, and needs its key.

The dashboard: totals, reliability, activity over time and calls by tool
Fig. 5An hour of calls on two public repositories: totals, reliability, activity and each tool.
One call opened: the question, the small model's steps turn by turn, and the reply
Fig. 6One ask opened: the question, the small model's steps turn by turn, and the reply Claude Code got, its citations checked.

05Organisations

On your machines, for the whole team.

Otōto runs where your code is, and talks only to the model servers you give it. Nothing is sent to us, and nothing checks for updates unless you run ototo update. For a company rolling it out to many machines:

Your model servers

vLLM, llama.cpp, Ollama, LM Studio or any OpenAI-compatible server. An organisation can make its own the only ones, so code goes nowhere else.

Settings you enforce

One file under every user's own; what it lists as enforced applies whatever a user sets: model servers, telemetry, trusted plugin signers. ototo doctor says what is enforced.

Managed rollout

ototo managed writes what to push with MDM, Ansible or a golden image: Claude Code's managed settings and MCP registration, Otōto's settings and trusted keys.

Usage, savings and the fleet

Counts, never code, over OpenTelemetry beside Claude Code's own, to a hub of collector, Prometheus and Grafana: usage and savings per person and team, and the versions and servers running now.

Plugins you trust

Only plugins signed by keys you trust load. An organisation can ship its own plugins and signers, and a forge plugin reaches only the hosts granted to it.

A security review

What runs, what it reads and writes, what leaves the machine, written down. Releases are signed and carry a CycloneDX SBOM; ototo update can read from your own mirror.

Rolling it out to a team? hello@ototo.dev

06Numbers

What it saved, measured.

41%
less Claude spend on our benchmark, every answer still rightnine coding tasks we wrote, three runs each
16%
less on 27 real bug fixes in four Java libraries, as many fixedjudged by each library's own tests, three runs each
38%
less than Claude Code's own Explore agent on the same fixessame tasks, same model

Claude Code on Claude Opus 5.5. Your numbers will differ with your code, your questions and your model server.

07Install

One command.

  1. curl -fsSL https://ototo.sh | sh

    For macOS on Apple silicon and Linux (x86_64, arm64). It downloads the newest release, checks it against Otōto's release key, and runs its installer, which finds your model server and sets up Claude Code and OpenCode, asking first. Name the server with sh -s -- --base-url http://gpu:8000/v1. Rather see it first? Read the script, then copy it or save it and run it yourself.

  2. ototo doctor

    Checks the settings, the model servers (with a tool-call test), Claude Code and OpenCode, and the plugins.

  3. ototo update

    When there is a newer release: checked the same way before anything changes. Nothing updates by itself.

08What's new

In the latest release.

2.8.6 beta

  • A tidier ototo --help. The commands are grouped by what they do: asking the small model (ask, locate, callers, edit, routine), reading the code with no model (read, outline, find, changes, history, replace, digest), watching and checking (tail, ui, doctor, report), setting up and updating (init, update, managed, settings), and the server Claude Code and OpenCode start (serve). Each has a one-line description; ototo <command> --help has the rest.

Coming from 2.8.4 or earlier? 2.8.5's changes are new to you too: routines, and the reporting names (INSTALL.md, "New in beta 2.8.5").

Every release

Open source, soon.

Otōto's source, its plugins and the kit for writing your own will be public under MIT or Apache-2.0, your choice, once the repository is ready for contributions. Until then, releases are signed builds.