Skip to content

Latest commit

 

History

1,065 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AFM — Your Mac is the cloud

AFM — local AI infrastructure for Apple Silicon

Swift 6.2+ macOS 26+ OpenAI compatible MIT

Website · Documentation · GitHub releases

What's new

AFM v0.9.17 — four new models with standard and MTP Qwen launch commands

AFM turns an Apple Silicon Mac into a private, OpenAI-compatible AI server. Run Hugging Face MLX models or Apple’s on-device Foundation Model, then connect the clients and SDKs you already use.

  • Native Swift executable—no Python runtime for serving
  • Local inference—no cloud account or API key
  • Chat, streaming, tools, structured output, reasoning, and logprobs
  • Vision OCR, speech, embeddings, and a built-in WebUI
  • Prefix caching, concurrent decode, speculative decoding, and metrics
  • Importable Swift packages for apps that need in-process inference

AFM is for Apple Silicon Macs running current macOS/Xcode toolchains. MLX model weights download from Hugging Face the first time you use them.

Install

Note

Stable v0.9.17 is the recommended release. It adds automatic Qwen 3.8 MTP sidecar discovery and quant-matched download behavior, plus expanded Qwen 3.8 tool-calling qualification. Install afm-next only to preview changes made after v0.9.17.

The qualified nightly and v0.9.17 are essentially the same build. Nightly nightly-20260816-bc343f6 was promoted to this stable release; the differences are release versioning and distribution packaging, not runtime functionality. Use the stable release unless a newer nightly explicitly lists post-v0.9.17 changes you need.

Stable (v0.9.17) Nightly (afm-next)
Homebrew brew install scouzi1966/afm/afm brew install scouzi1966/afm/afm-next
pip pip install macafm pip install --extra-index-url https://maclocal-ai.pages.dev/afm/wheels/simple/ macafm-next
Release notes v0.9.17 Latest nightly

Install a previous version

Older stable releases are kept as pinned formulae in the Homebrew tap and as version-pinned wheels on PyPI. This is useful for reproducing an issue against a specific build or rolling back without waiting for a new release.

Homebrew (pinned stable formulae): afm@<version> — available for 0.9.0, 0.9.1, and 0.9.30.9.10.

brew install scouzi1966/afm/afm@0.9.10
brew uninstall afm
brew link afm@0.9.10
afm --version

Homebrew (pinned nightly formulae): afm-next@<full-version> — for example, afm-next@0.9.15-next.20260808.e70cc52. See the Homebrew tap for available pinned nightlies.

brew install scouzi1966/afm/afm-next@0.9.15-next.20260808.e70cc52

pip (version-pinned wheels): install any published release by version.

pip install macafm==0.9.10
pip install --extra-index-url https://maclocal-ai.pages.dev/afm/wheels/simple/ \
  macafm-next==0.9.15.dev20260808

Start in two minutes

brew install scouzi1966/afm/afm

# Start a small MLX model and open the WebUI
afm mlx -m Qwen3-0.6B-4bit -w

AFM is now listening at http://127.0.0.1:9999/v1.

curl http://127.0.0.1:9999/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen3-0.6B-4bit",
    "messages": [{"role": "user", "content": "Explain unified memory in one paragraph."}],
    "stream": false
  }'

Or use Apple’s on-device model:

afm -w

Native terminal chat

Use --tui when you want a private, full-screen chat without running an HTTP server:

# Apple Foundation Models
afm --tui

# Any supported MLX model (all normal sampling/runtime flags still apply)
afm mlx -m Qwen3-0.6B-4bit --tui

TUI changes have a model-free regression harness. make test-tui runs stable Markdown/math/code snapshots and exercises keyboard input, terminal sizing, alternate-screen cleanup, and raw-mode restoration through a real macOS pseudo-terminal. If an intentional visual change updates the expected output, run Scripts/test-tui.sh --record, inspect the snapshot diff, then rerun make test-tui normally. The same focused suite runs automatically on relevant pull requests.

The terminal UI streams responses, separates optional reasoning, and renders both answers and visible reasoning through a native CommonMark/GFM renderer. Headings, nested/task lists, quotes, tables, links, inline formatting, fenced code, and raw HTML are presented as inert terminal output. Code uses source-compiled Tree-sitter grammars for semantic highlighting and line numbers (details); unified diffs have distinct file, hunk, addition, and deletion styling. Inline and display LaTeX are rendered as readable Unicode math, including fractions, roots, super/subscripts, operators, Greek symbols, matrices, and cases. The UI also reports token and throughput statistics.

Reasoning is collapsed by default into a live activity row with its phase, animated cursor, elapsed time, and generated character count. Press Tab during generation to expand or collapse the reasoning panel without interrupting the model. Use /reasoning expanded, /reasoning collapsed, /reasoning off, or /reasoning last to control it explicitly.

It supports multiline editing, prompt history, cancellation, persisted/searchable sessions under ~/.afm/sessions, transcript export, attachments, terminal-width-aware tables, themes, and safe actions for response artifacts. Use /help for the command palette.

The navigation follows Codex CLI conventions. Normal chat output remains in Terminal scrollback. Press Ctrl+T to open the full transcript overlay, then scroll with a Mac trackpad or mouse wheel, arrows, Page Up/Down, Ctrl-U/Ctrl-D, or Home/End; press Ctrl+T again to close it. Add --no-alt-screen to keep overlays inline too. /blocks opens a session-wide code-block list: navigate with arrows or paging keys, press Enter for Copy/Save/Preview actions, and Escape to return. Numbered /save, /copy, and /open commands remain available for direct use.

Code is never executed automatically. /save refuses overwrites unless /save! is used, and only an explicit /open previews HTML or JavaScript in the browser. iTerm2 and Kitty can display local images inline; Terminal.app uses an explicit /image Quick Look fallback.

Choose your runtime

Runtime Best for Start it
MLX Open models, VLMs, agent controls, performance tuning afm mlx -m <model>
Apple Foundation Models Zero-download system model and .fmadapter LoRA adapters afm
DwarfStar Compatible fixed-schedule Metal checkpoints afm mlx -m <owner/repo> (auto-resolved) or afm mlx -m <checkpoint.gguf> --mlx-runtime dwarfstar
Gateway One model list for Ollama, LM Studio, Jan, and other local servers afm --gateway

Model IDs without an organization default to mlx-community, so Qwen3-0.6B-4bit and mlx-community/Qwen3-0.6B-4bit both work.

Evaluate a local model

AFM ships all 91 labeled variants from the repository's comprehensive MLX test as a deterministic, no-judge suite. The model loads once, every output and timing measurement is retained locally, and a self-contained HTML report opens when the run finishes.

afm mlx -m mlx-community/Qwen3-0.6B-4bit --eval

# Headless run, suite discovery, and custom-suite scaffolding
afm mlx -m <model> --eval --no-open
afm mlx --eval-list
afm mlx --eval-init my-suite
afm mlx --eval-validate ~/.afm/evals/my-suite.json
afm mlx -m <model> --eval-suite comprehensive --eval-suite my-suite

Run artifacts are stored in collision-safe ~/.afm/evals/<date-time>-<model>-<suite>/ directories. See Local model evaluations for the suite schema, deterministic checks, report contents, and security limits.

Why AFM works well for agents

AFM is built for multi-turn, tool-using clients—not only chat demos.

Capability What it gives you
Native tool formats Auto-detection for JSON, Qwen XML, Gemma, GLM, Kimi, MiniMax, LFM2, and related formats
Tool choice auto, none, required, and named-function forcing
Streaming tool deltas OpenAI-style tool-call chunks while ordinary content continues to stream
Structured output json_object, json_schema, and token-level xgrammar enforcement when enabled
Reasoning extraction <think> and harmony analysis channels mapped to reasoning_content
Determinism and inspection seed, logprobs, top_logprobs, request IDs, tracing, and raw-parser mode
Long-running reliability Cancellation, Retry-After, token counting, fair concurrent queues, and Prometheus metrics
Prefix reuse Radix-tree KV caching for stable system prompts and multi-turn agent loops

Pick a tool-calling mode

  • Native (default): AFM detects the model’s own format and uses the narrowest parser. Use this for parity checks and benchmarks.
  • Repair: add --tool-call-parser afm_adaptive_xml for JSON-in-XML fallback, type coercion, nullable-schema handling, and fuzzy tool-name matching. Add --fix-tool-args when a model renames arguments.
  • Raw: add --tool-call-parser none to return the model’s tool markup as ordinary assistant content.

See MLX tool-calling modes for examples and benchmark guidance.

Connect an existing client

Most OpenAI-compatible clients need only a base URL and a placeholder API key:

Base URL: http://127.0.0.1:9999/v1
API key:  x

Copy-ready guides:

OpenCode · OpenClaw · Cline · Continue · Aider · Cursor · Hermes

OpenClaw users can also generate a provider block directly:

afm mlx -m Qwen3-Coder-Next-4bit --openclaw-config

API surface

Method Endpoint Purpose
POST /v1/chat/completions Chat, SSE streaming, tools, reasoning, structured output, logprobs
GET /v1/models Active model and gateway model discovery
POST /v1/embeddings Apple NaturalLanguage embeddings for RAG and semantic search
POST /v1/vision/ocr OCR, tables, barcodes, classification, saliency, and PDFs
POST /v1/audio/transcriptions On-device speech-to-text
POST /v1/audio/speech Text-to-speech using installed Apple voices
POST /v1/tokenize vLLM-compatible tokens and counts for the loaded MLX model
POST /v1/count_tokens Anthropic-style input token count
POST /v1/batch/completions Multiplex up to 64 completions over SSE
POST /v1/chat/completions/{id}/cancel Cancel an in-flight generation
GET /metrics Prometheus queue, token, throughput, and timing metrics
GET /openapi.json OpenAPI description
GET /docs Interactive API reference served by AFM

AFM also implements OpenAI-style file and batch-job endpoints under /v1/files and /v1/batches when the MLX batch service is active.

Apple-native tools

The CLI and HTTP server expose useful system frameworks without another service.

# OCR text or a table from an image/PDF
afm vision --file invoice.pdf --table

# Other Vision modes: text, table, barcode, classify, saliency, auto
afm vision --file photo.heic --mode classify --format json

# Speech recognition
afm speech transcribe --file meeting.wav --format srt

# Text to speech
afm speech synthesize "Hello from AFM" --voice nova --output hello.aac

# Dedicated OpenAI-compatible embeddings server (default port 9998)
afm embed

For vision-language models, add --vlm and pass one or more files with --media.

Performance controls

Defaults are a good starting point. Use these when the workload calls for them:

# Reuse prompt KV across requests
afm mlx -m <model> --enable-prefix-caching

# Save memory on long context
afm mlx -m <model> --kv-bits 8

# Fair-queue concurrent requests through one model
afm mlx -m <model> --concurrent 4

# Strict tool/JSON schemas with xgrammar
afm mlx -m <model> --enable-grammar-constraints

# Per-request device, memory, timing, and bandwidth estimates
afm mlx -m <model> --gpu-profile -s "Explain Metal kernels"

Supported checkpoints can also use speculative decoding:

  • --mtp for compatible Qwen models. Qwen3.8 automatically prefetches the separately published MTP head matching the base checkpoint's quantization; use --mtp-model <repo-or-path> to override it.
  • --eagle3 <drafter-directory> for supported dense Gemma4 models
  • --dspark-support <support.gguf> for compatible DwarfStar DSpark workflows

Read decode optimizations before choosing a checkpoint or interpreting benchmark results.

Sampling and response controls

The MLX backend supports temperature, top_p, top_k, min_p, repetition_penalty, presence_penalty, seed, stop, logprobs, and top_logprobs.

Useful server defaults:

# Apply one JSON schema when requests omit response_format
afm mlx -m <model> \
  --guided-json '{"type":"object","properties":{"answer":{"type":"string"}},"required":["answer"]}' \
  --enable-grammar-constraints

# Disable model reasoning/thinking
afm mlx -m <model> --no-thinking

# Pin chat-template keyword arguments
afm mlx -m <model> --chat-template-kwargs '{"enable_thinking":false}'

Use AFM as a Swift package

The repository publishes focused Swift Package Manager products:

  • AFMKitCore — provider contracts and core types
  • AFMOpenAICompat — OpenAI-compatible request/response types
  • AFMKitMLX — MLX model loading and inference
  • AFMKitFoundationModels — Apple Foundation Models backend
  • AFMKitFoundationModels27 — macOS 27 provider protocol adapters
  • AFMKitFoundationModels27DwarfStar — opt-in DwarfStar macOS 27 adapter
  • AFMKitDwarfStar — DwarfStar runtime integration
  • AFMKitServices — vision, speech, and embedding services
  • AFMKit — high-level headless inference facade
  • AFMServer — Vapor HTTP layer
  • afm — CLI executable
dependencies: [
    .package(
        url: "https://github.com/scouzi1966/maclocal-api.git",
        branch: "main"
    )
]

Start with the AFMKit public API guide and the consumer examples.

Build from source

git clone https://github.com/scouzi1966/maclocal-api.git
cd maclocal-api
./build.sh

The complete build initializes submodules, applies AFM-owned vendor patches, builds the WebUI, rebuilds Metal resources when the toolchain is available, and creates the release executable. Add --install to install it on your PATH.

Requirements

  • Apple Silicon Mac
  • macOS 26 or newer for the complete feature set
  • Xcode 27 for development builds
  • Disk and unified memory appropriate for the model you choose

Small 0.6B–4B quantized models are the easiest way to confirm a setup. Large 30B-class models need substantially more unified memory.

Documentation map

Contributing

Issues, reproducible test cases, documentation improvements, and model-compatibility reports are welcome. Read AGENTS.md and CLAUDE.md before changing build, test, or vendored integration code.

If AFM is useful to you, star the repository. You may also like Vesta AI Explorer, a full-featured native macOS AI app.

License

MIT

About

'afm' command cli: macOS server and single prompt mode that exposes Apple's Foundation and MLX Models and other APIs running on your Mac through a single aggregated OpenAI-compatible API endpoint. Supports Apple Vision and single command (non-server) inference with piping as well . Now with Web Browser and local AI API aggregator

Topics

Resources

Stars

330 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages