| Local LLM Testing & Benchmarking for Apple Silicon | Community Leaderboard |
brew install --cask uncsoft/anubis/anubis-oss
Or download the zip directly from the Releases page and drag to /Applications.
π¨ Benchmark analysis is live! Check out the results here, over 400+ community submitted runs analyzed Benchmark Report
Anubis is a native macOS app for benchmarking, comparing, and managing local large language models using any OpenAI-compatible endpoint - Ollama, MLX, oMLX, LM Studio Server, OpenWebUI, Docker Models, etc. Built with SwiftUI for Apple Silicon, it provides real-time hardware telemetry correlated with full, history-saved inference performance - something no CLI tool or chat wrapper offers. Export benchmarks directly without having to screenshot, and export the raw data as .MD or .CSV from the history. You can even OLLAMA PULL models directly within the app.
New in 3.7: the Flow Builder β drag-and-drop sequencer for multi-model / multi-prompt / N-rep benchmark runs, with share-ready 16:9 or 1:1 report cards. Build it once, run it hands-off, post the PNG.
Build, save, and share repeatable benchmark recipes β a Shortcuts-style editor where you sequence steps like Set Backend β Set Model β Repeat Γ 5 β Run Benchmark β Unload and play them back hands-off. Every individual run still lands in normal Run History, and the whole sequence gets a single share-ready report.
Why use a flow instead of clicking Run repeatedly?
.anubisflow JSON to share, version, or move between machinesEditor. Three panes: a left Palette of step blocks, a center Step List rendered as a nestable tree, a right Inspector that swaps controls per step type. Add steps by clicking a palette block or dragging it into precise position. Reorder by drag, by β/β chevrons, or by editing the tree directly. Containers (Repeat, For Each Model, For Each Prompt) hold child steps with an indented rail.
Live lint. As you build, Anubis flags dangling Set steps (e.g. a Set Model with no following Run Benchmark), empty For Each lists, blank entries, and Run Benchmarks that are missing required prior Sets. The header chip jumps straight to the offending step.
For Each, explicitly. Two rules worth remembering:
For Each Model, the active model is whatever the iterator just picked β any prior Set Model is ignored for that iteration (the lint catches this).3 models Γ 2 prompts Γ 5 reps = 30 BenchmarkSessions.Templates. Sidebar + β New from Template ships six starters: Quick Smoke Test, Repeated Run (5Γ), Cold vs Warm Start, Multi-Model Comparison, Q4 vs Q8 Quantization, Prompt Variants. Each is a real working flow; edit or use as-is.
Run sheet. Live progress with checkmarks per step, a scrollback log, a βRun X of Yβ counter, and a per-run latest-session preview. Stop (β.) cancels cleanly mid-run.
Reports built for sharing. Post-run, View Report (β§βR) opens a forced-dark, 1920Γ1080 card with a hero, winner callout, per-model leaderboard with relative bars, per-rep tables, and a methodology footer. Toggle to a 1080Γ1080 square for Instagram/X. Save as PNG @ 2Γ retina (3840Γ2160 or 2160Γ2160), copy to clipboard, or save the per-model CSV. Past runs live in the sidebarβs Run History section.
Import / export. Right-click any flow β Exportβ¦ writes a portable .anubisflow JSON file. + β Import .anubisflowβ¦ reads it back. Format versioned so older clients refuse newer files instead of silently misreading them.
Open the in-app Help sheet (the ? in the Flows sidebar header) for the full reference, including worked code-style examples of each For Each pattern.
The History window is now genuinely usable, and discovering Ollama models no longer requires leaving the app.
https://ollama.com/library directly. Cards show description, capability chips, size buttons, pull count, last-updated; click a size to pull through the existing progress UI. Search hits /search?q=β¦. Already-installed models show a green badge. 24-hour client-side cache, Refresh button to bypass it.think field, and what to do when a model rejects it.Bundle N consecutive runs of the same configuration into one benchmark with mean and 95% bootstrap confidence intervals, plus a significant correction to system-wide power accounting.
run_group_id, sample count, rep index, seed strategy, and mean Β± CI for the headline metrics. The leaderboard page renders βΒ±CI Β· N repsβ inline under tok/s on group rows; the explorer surfaces the group aggregates as sortable columns.system_power and the headline J/Tok metric was a corresponding undercount. Per-channel scaling is now correct; methodology version tag bumps 1 β 2 so cross-version comparisons stay honest. Migration v7 in the local app DB retroactively rescales historical sessions.Connection: close header weβd added for Ollama chunk pacing was causing LM Studio to reject the immediately-following request). Auto-detection now correctly identifies the inference worker at ~/.lmstudio/.internal/utils/node, with a self-healing soft pin that re-evaluates every 2 s. Stops mis-attributing to Ollama after a model switch..userInteractive priority; @Published cascades batched at 1 Hz during multi-rep group streaming; font-level ligature disable bypasses the CopyOfFontWithLigatureSetting hot path that caused the 25-rep hang.Tri-state control over Ollamaβs think request parameter, exposed in the Benchmark Performance disclosure when the Ollama backend is selected.
think:true to enable reasoning where supportedthink:false to disable reasoning on models that default it on (e.g. recent DeepSeek-R1 builds)The choice persists across launches.
Output tokens/sec is now visible-throughput only for reasoning models. Previously, thinking time was charged against TTFT and thinking tokens were counted as output, inflating the numbers. Fixes #17 and #18.
reasoning_content, reasoning, or inline <think>β¦</think> tags) and surface it wrapped in <think>β¦</think> markers in the responseAnubis now benchmarks Appleβs on-device Foundation Model alongside Ollama, MLX, and the rest β no server, no network, no setup. If your Mac supports Apple Intelligence (macOS 26+), it shows up in the backend menu automatically.
Apple Intelligence from the backend selector and run; it talks directly to the on-device model via Appleβs FoundationModels frameworkInstructionsExport the per-model performance summary directly from the Reports tab.
Push your Apple Silicon to its limits and observe power draw, thermal throttling, and frequency scaling under controlled load - all from within the Monitor.
yes processes per core. Choose All Cores, P-Cores only, E-Cores only, or Single Corememcpy to saturate the memory bus. Reports measured bandwidth in GB/s, directly comparable to your chipβs theoretical max. Three pressure levels (Light 25% / Moderate 50% / Heavy 75% of free memory)A compact, frameless, always-on-top overlay showing live system metrics - launchable from any tab via the sidebar or from the Monitorβs Float button.
Five new built-in prompts covering causal reasoning, system design, dialogue writing, historical analysis, and constrained writing - bringing the total to 15 across five categories.
The local LLM ecosystem on macOS is fragmented:
Anubis fills that gap - all in a native macOS app.
Real-time performance dashboard for single-model testing.
/v1 suffix from backend URLs to prevent double-pathing errorsSide-by-side A/B model comparison with the same prompt.
Drag-and-drop sequencer for repeatable, multi-step benchmark recipes.
2 models Γ 3 prompts Γ 5 reps = 30 runs shown live in the header chip.anubisflow import / export β versioned JSON, portable between machines? in the sidebar header) covers the editor, every step type, and For Each rules with worked examplesStandalone real-time hardware monitoring dashboard - no benchmark required.
Upload your benchmark results to the community leaderboard and see how your Mac stacks up against other Apple Silicon machines.
Unified model management across all backends.
~/.lmstudio/models/ and ~/.cache/huggingface/hub/ for disk size, quantization, and pathAnubis checks for updates automatically via Sparkle and notifies you when a new version is available.
Reports and Exports
GPU Core detail
Arena Mode
Settings (add connections with quick presets)
Vault - View model details, unload, and Pull models directly for Ollama
| Backend | Type | Default Port | Setup |
|---|---|---|---|
| Apple Intelligence | On-device (Foundation Models) | β | macOS 26+ with Apple Intelligence enabled. No setup; appears in the backend menu when supported. |
| Ollama | Native support | 11434 | Install from ollama.com - auto-detected on launch |
| LM Studio | OpenAI-compatible | 1234 | Enable local server in LM Studio settings |
| mlx-lm | OpenAI-compatible | 8080 | pip install mlx-lm && mlx_lm.server --model <model> |
| vLLM | OpenAI-compatible | 8000 | Add in Settings |
| LocalAI | OpenAI-compatible | 8080 | Add in Settings |
| Docker ModelRunner | OpenAI-compatible | user selected | Add in Settings |
Any OpenAI-compatible server can be added through Settings > Add OpenAI-Compatible Server with a name, URL, and optional API key.
Anubis captures Apple Silicon telemetry during inference via IOReport and system APIs:
| Metric | Source | Description |
|---|---|---|
| GPU Utilization | IOReport | GPU active residency percentage |
| CPU Utilization | host_processor_info |
Usage across all cores |
| GPU Power | IOReport Energy Model | GPU power consumption in watts |
| CPU Power | IOReport Energy Model | CPU (E-cores + P-cores) power in watts |
| ANE Power | IOReport Energy Model | Neural Engine power consumption |
| DRAM Power | IOReport Energy Model | Memory subsystem power |
| GPU Frequency | IOReport GPU Stats | Weighted average from P-state residency |
| Process Memory | proc_pid_rusage |
Backend process phys_footprint (includes Metal/GPU allocations) |
| Thermal State | ProcessInfo.thermalState |
System thermal pressure level |
Anubis automatically detects which process is serving your model:
lsof to find the PID listening on the inference port (called once per benchmark start)phys_footprint (same as Activity Monitor) which includes Metal/GPU buffer allocations - critical for MLX and other GPU-accelerated backendsMetrics degrade gracefully - if IOReport access is unavailable (e.g., in a VM), Anubis still shows inference-derived metrics.
# macOS - install Ollama
brew install ollama
# Start the server
ollama serve
# Pull a model
ollama pull llama3.2:3b
The easiest path is Homebrew β installs the signed, notarized .app from the latest GitHub release:
brew install --cask uncsoft/anubis/anubis-oss
Or download the zip directly from the Releases page and drag to /Applications.
Anubis auto-updates via Sparkle in either case. Once installed, Anubis will auto-detect Ollama on launch. Other backends can be added in Settings.
The app updates itself via Sparkle β just open Anubis and accept the update prompt (or Settings β Check for Updates). Thatβs the recommended path and works for every install method.
To update through Homebrew instead:
brew update
brew upgrade --cask anubis-oss --greedy
The
--greedyflag is required because the cask is markedauto_updates(the app updates itself), so a plainbrew upgradeintentionally skips it.
Cask 'anubis-oss' is unreadable: syntax errors / shows <<<<<<< conflict markersgit clone https://github.com/uncSoft/anubis-oss.git
cd anubis-oss/anubis
open anubis.xcodeproj
In Xcode:
Cmd+R)After a benchmark completes, click the Upload button in the benchmark toolbar to submit your results to the community leaderboard. Enter a display name and your run will appear in the rankings - no account required. Only performance metrics and hardware info are submitted; response text is never uploaded.
# Clone
git clone https://github.com/uncSoft/anubis-oss.git
cd anubis-oss/anubis
# Build via command line
xcodebuild -scheme anubis-oss -configuration Debug build
# Run tests
xcodebuild -scheme anubis-oss -configuration Debug test
# Or just open in Xcode
open anubis.xcodeproj
Resolved automatically by Swift Package Manager on first build:
| Package | Purpose | License |
|---|---|---|
| GRDB.swift | SQLite database | MIT |
| Sparkle | Auto-update framework | MIT |
| Swift Charts | Data visualization | Apple |
Anubis follows MVVM with a layered service architecture:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β PRESENTATION LAYER β
β BenchmarkView ArenaView MonitorView VaultView Settings β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β SERVICE LAYER β
β MetricsService InferenceService ModelService Export β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β INTEGRATION LAYER β
β OllamaClient OpenAICompatibleClient IOReportBridge ProcessMonitor β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β PERSISTENCE LAYER β
β SQLite (GRDB) File System β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Views display data and delegate to ViewModels. ViewModels coordinate Services. Services are stateless and use async/await. Integrations are thin adapters wrapping external systems (Ollama API, IOReport, etc.).
anubis/
βββ App/ # Entry point, app state, navigation
βββ Features/
β βββ Benchmark/ # Performance dashboard
β βββ Arena/ # A/B model comparison
β βββ Flows/ # Flow Builder: editor, executor, report, lint, templates
β βββ Monitor/ # System monitor, stress tests, floating HUD
β βββ Reports/ # Cross-run model performance table
β βββ Vault/ # Model management
β βββ Settings/ # Backend config, about, help, contact
βββ Services/ # MetricsService, InferenceService, FlowExecutor, ExportService
βββ Integrations/ # OllamaClient, OpenAICompatibleClient, IOReportBridge, ProcessMonitor
βββ Models/ # Data models (BenchmarkSession, Flow, ModelInfo, etc.)
βββ Database/ # GRDB setup & migrations
βββ DesignSystem/ # Theme, colors, reusable components
βββ Demo/ # Demo mode for App Store review
βββ Utilities/ # Formatters, constants, logger, benchmark prompts
All inference backends implement a shared protocol, making it straightforward to add new ones:
protocol InferenceBackend {
var id: String { get }
var displayName: String { get }
var isAvailable: Bool { get async }
func listModels() async throws -> [ModelInfo]
func generate(prompt: String, parameters: GenerationParameters)
-> AsyncThrowingStream<InferenceChunk, Error>
}
All data is stored locally - nothing leaves your machine.
| Data | Location |
|---|---|
| Database | ~/Library/Application Support/Anubis/anubis.db |
| Exports | Generated on demand (CSV, Markdown) |
| Preferences | UserDefaults |
# Make sure Ollama is running
ollama serve
# Verify it's accessible
curl http://localhost:11434/api/tags
ollama pull <model-name>Contributions are welcome. A few guidelines:
errorDescription and recoverySuggestionIntegrations/ implementing InferenceBackendInferenceServiceSettings/If Anubis is useful to you, consider buying me a coffee on Ko-fi or sponsoring on GitHub. It helps fund continued development and new features.
A sandboxed, less feature rich version is also available on the Mac App Store if you prefer a managed install.
GPL-3.0 License - see LICENSE for details.