Skip to content

Before you start ​

Install only what you plan to use. Start with AgentX and one working agent. Add voice, desktop control, or a second machine later.

Choose your starting point ​

I want to…What I needStart here
Explore without a model accountA source installation; the demo uses scripted repliesSource setup, then demo
Run agents and use the browser dashboardDocker or Node.js; a model connectionCore setup
Chat inside the dashboardRunning dashboard, daemon, and configured agentIn-page chat
Use the OpenCode terminal interfaceCore setup + OpenCode v2 or newerTerminal setup
Talk to agents on my MacCore setup + compatible Mac + speech setupDesktop setup
Point at controls or inspect my screenDesktop helper + permissions + relevant model credentialsComputer use
Connect agents on two machinesCore setup on both + private network + pairing the machinesNetworking

The daemon is the background service that runs agents. The dashboard is the browser interface connected to it. A provider supplies the AI model. An API key is a private credential for that provider. The examples below use an agent with the ID support; use your own agent's ID.

Core setup ​

Install Git if running git --version in a terminal says the command is not found. The install instructions use Git to download AgentX.

Choose one installation path. Docker runs the daemon and dashboard in containers. A source installation runs them directly on your machine.

Option A: use Docker ​

  1. Install Docker Desktop or Docker Engine. Start Docker before continuing.

  2. Terminal: check that Docker answers:

    sh
    docker version
    docker compose version
  3. Follow Install AgentX with Docker.

Ready when: docker version shows a responding server and the dashboard opens after starting Compose. You do not need host Node.js for this path. Native macOS desktop tools must be installed on the Mac itself; they do not run inside the Linux container.

Option B: run from source ​

  1. Install Node.js 22.19 or newer, up to 26, from Node.js downloads. Node 22 is what the maintainers run day to day.

  2. Terminal: install pnpm 10 (the package manager AgentX is built with) and check both versions:

    sh
    npm install -g pnpm@10
    node --version
    pnpm --version
  3. Confirm Node reports v22.… and pnpm reports 10.…, then follow Run from source.

Help: pnpm installation. If installation reports a native compilation error, follow node-gyp's platform prerequisites for Python and a C/C++ toolchain.

Commands in these guides

agentx means an installed CLI. When running from source, use node dist/cli.js in its place after pnpm build. Run commands from the directory containing your agentx.json. The install page explains the 0.27.0 package limitation.

Connect one model ​

For a first setup, the browser wizard's Anthropic API (BYO key) option is a direct route:

  1. Follow Anthropic's getting-started guide to obtain API access and a key.
  2. Browser: open AgentX /setup, select that engine, and enter the key.
  3. Save the agent and restart the daemon as described in Your first agent.
  4. Send a short test message.

Ready when: the agent returns a real response. Other engines need their own installed runtime and authentication. Installing OpenCode, Docker, or the desktop app does not provide model credentials. Provider usage and account requirements are separate from AgentX; see costs.

Desktop assistant ​

Required: Apple Silicon Mac, macOS 14 or newer, Node.js 22, configured AgentX agent, and Apple's command-line tools. Check Apple menu → About This Mac for your chip and macOS version.

  1. Terminal: install Apple's tools if needed:

    sh
    xcode-select --install

    Complete the macOS installer, then check xcrun --find swiftc. See Apple's developer tools help.

  2. Terminal: from the folder that holds your agentx.json, install and start the assistant:

    sh
    agentx desktop install --agent support

    Replace support with an ID from agentx agent list.

  3. Complete speech setup if you want voice, and permissions for the features you use.

  4. Terminal: run agentx desktop status to check that the assistant is set to start at login.

The installer handles: building both native apps, installing them under ~/Applications, remembering your agent and daemon URL, and starting the assistant at login.

You handle: installing Apple's tools, configuring the daemon/model, setting up a speech backend, and accepting macOS permissions. The installer does not download Whisper or grant permissions.

Voice input and spoken replies ​

Choose one transcription option. You can later configure both for fallback.

Option A: ElevenLabs ​

  1. Create an API key following ElevenLabs' key guide. Enable the access needed for speech-to-text and, if used, text-to-speech.

  2. Terminal: create the local configuration folder and open a key file:

    sh
    mkdir -p ~/.agentx
    touch ~/.agentx/elevenlabs-key.txt
    chmod 600 ~/.agentx/elevenlabs-key.txt
    nano ~/.agentx/elevenlabs-key.txt
  3. Paste only the key, save with Control–O, press Enter, and exit with Control–X. A key file works for login-launched apps without relying on terminal environment variables.

  4. Terminal: restart the assistant:

    sh
    agentx desktop stop
    agentx desktop start

Ready when: holding Option–Space, speaking a short question, and releasing produces a transcript and reply. This option needs internet access and available provider quota. Official help: ElevenLabs quickstart.

Option B: local Whisper ​

This is an optional local transcription setup for Apple Silicon. It needs Python, FFmpeg, and the mlx-whisper package. The MLX Whisper project documents installation and model downloads.

  1. Terminal: install Python and FFmpeg (an audio converter). If you use Homebrew:

    sh
    brew install python ffmpeg
  2. Terminal: install Whisper into its own Python environment (a venv: a private folder of Python packages that doesn't affect the rest of your Mac):

    sh
    python3 -m venv ~/.agentx/whisper-env
    ~/.agentx/whisper-env/bin/python -m pip install mlx-whisper
  3. Terminal: create a launcher at the path AgentX expects. It also makes Homebrew's FFmpeg available to the login app:

    sh
    mkdir -p ~/.local/bin
    cat > ~/.local/bin/mlx_whisper <<'SH'
    #!/bin/sh
    export PATH="/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:$PATH"
    exec "$HOME/.agentx/whisper-env/bin/mlx_whisper" "$@"
    SH
    chmod +x ~/.local/bin/mlx_whisper
    ~/.local/bin/mlx_whisper --help

    If you already have a launcher there, keep it and check that it works instead of replacing it.

  4. Terminal: test with a short audio file of your own:

    sh
    ~/.local/bin/mlx_whisper /path/to/test.wav --model mlx-community/whisper-large-v3-turbo

    Replace the audio path. The first run downloads the model and needs internet access and free disk space. Let it finish before testing the desktop hotkey.

  5. In the terminal, run the desktop install again so the assistant can find FFmpeg when it starts at login:

    sh
    agentx desktop install --agent support

    A login app does not see your terminal's settings. The installer records the folder where it found FFmpeg. If it prints ffmpeg was not found, finish step 1 and run it again.

  6. In the terminal, run agentx doctor. Under Desktop, it should report ffmpeg reachable by the desktop app.

Ready when: the test creates a text transcript. Restart the desktop assistant and test Option–Space. Without ElevenLabs, spoken replies use macOS say. Local transcription does not make a remotely hosted agent model work offline.

Option C: Parakeet, built into the app ​

The desktop app can also transcribe with Parakeet, which needs no Python and no FFmpeg. It downloads a 483 MB model the first time and has no Arabic. Keep Option B installed as well: Whisper answers while Parakeet downloads or loads. If you speak Arabic, stay on Whisper. See Speech to text on this Mac to switch it on.

macOS permissions ​

  1. Open System Settings → Privacy & Security.
  2. Select the permission from the table below.
  3. Turn on access for the app named in the macOS prompt.
  4. Relaunch the app if macOS asks.

Privacy & Security › Accessibility with access turned on for AgentX, which macOS lists as node

PermissionNeeded forApp normally requesting it
MicrophoneRecording your spoken requestAgentX Desktop
AccessibilityReading controls and interacting with appsAgentX Helper / the invoking app
Screen Recording or Screen & System Audio RecordingReading text off the screen (OCR, optical character recognition) and screen capturesAgentX Helper / the invoking app

Official help: Accessibility access, microphone access, screen recording access. An app may appear only after it first requests access. A running desktop login service alone does not prove these permissions are enabled.

Computer use ​

Install the desktop helper and grant the permissions above first. Then choose the capability:

CapabilityAdditional requirementCheck
Locate and highlight a control (point)Available ui-element decision seat (see below)agentx point "the search field" while that field is visible
Describe the screen (look)OpenRouter key and access to a vision modelagentx look "What window is visible?"
Verify a visible resultVision setup; optional screen-state seatagentx look "The command palette is open" --verify
Run a guided lessonDependencies needed by that lessonagentx teach lists the available lessons

For look, create a key using the OpenRouter quickstart, then save it as a single line in ~/.agentx/openrouter-key.txt with owner-only permissions, following the key-file steps above. Screen captures are sent to the vision provider. Test using a demo workspace.

Some features hand a small, fixed-choice question to a fast "decision" model instead of the main agent. Each such question is a seat (for example ui-element: which control on screen matches your words). Jev is the decision model AgentX uses for seats. To set it up, follow the decision-seat setup. The current jev backend reads OPENROUTER_API_KEY; the direct typesafe backend reads TYPESAFE_API_KEY. Put the appropriate variable in your AgentX .env, enable the desired seats, and restart the daemon. The vision key file is not a substitute for the decision backend's environment variable. Confirm backend access before enabling a seat in active mode.

Ready when: the specific command you intend to use works. point highlights without clicking; lessons can click and type. The computer-use walkthrough explains how to check results.

Terminal interface ​

  1. Install OpenCode (a terminal-based coding assistant) using its official v2 instructions.
  2. Terminal: run opencode --version. AgentX's integration requires v2 or newer.
  3. Terminal: start your daemon.
  4. Terminal: run agentx tui --agent support.

Ready when: OpenCode opens with the AgentX agent selected. If OpenCode is missing or too old, AgentX uses its built-in TUI. agentx tui --legacy selects that fallback directly. See TUI connection limits before using an authenticated remote daemon.

Two machines and A2A ​

A2A (agent-to-agent) is how an agent on one machine hands work to an agent on another. The connected machines form a mesh.

  1. Install AgentX and configure a model on each machine that will run agents.
  2. Install Tailscale using the official setup guide, and join your private network (Tailscale calls it a tailnet).
  3. Follow AgentX's Tailscale guide to configure reachable addresses, pair peers, and load matching tokens.
  4. Terminal: run agentx mesh list. Each peer shows healthy.
  5. Send a small task to the other machine.

Ready when: the remote agent answers. Tailscale is one private-network option; it is not required for a single-machine setup. The separate standalone A2A server needs its own provider runtime and authentication; it does not automatically use a configured daemon agent.

Channels and workflows ​

For Telegram, you need a Telegram account and bot token; follow Connect Telegram. Other channels need their own account, credentials, and event-delivery configuration; use Channels to choose a supported integration. Test one incoming message before scheduling work that depends on it.

For the workflow assistant, you need a configured agent and workflows turned on. Terminal:

sh
agentx config set workflows.enabled true

Follow the command's reload/restart result, then use Describe an automation. Ready when: the assistant returns a proposal and you can inspect it before applying or running it. A webhook or scheduled workflow also needs the corresponding channel or schedule configured.

Recording a tutorial ​

Screen Studio is optional. For the Screen Studio example, install the app and follow its recording guide, including macOS capture permissions. Confirm that a playable clip was saved.

agentx teach --record uses macOS screencapture instead. It does not require Screen Studio. See record a VS Code walkthrough.

Check it worked ​

  • [ ] My chosen installation path works.
  • [ ] My daemon is running and one agent answers a short message.
  • [ ] I installed only the optional features I want.
  • [ ] Each optional feature passes its own “Ready when” check.

There is no measured universal RAM or disk minimum for AgentX yet. Local speech models, several agents running at once, and kept recordings all change what you need.

If something is wrong ​

  • docker version shows no server: Docker Desktop isn't running. Start it and try again.
  • node --version is below v22.19 or above v26: install Node.js 22. AgentX doesn't run on other versions.
  • Installing reports a native compilation error: install the node-gyp platform prerequisites, then install again.
  • An app is missing from System Settings › Privacy & Security: use the feature once so the app asks for access, then look again.
  • ffmpeg was not found during agentx desktop install: finish the FFmpeg step under local Whisper, then run the install again.
  • Anything else: run agentx doctor and follow It's not answering.

Released under the MIT License.