diff --git a/docs/wiki/Agent-CLIs.md b/docs/wiki/Agent-CLIs.md index eb013b29..d7ad4b22 100644 --- a/docs/wiki/Agent-CLIs.md +++ b/docs/wiki/Agent-CLIs.md @@ -273,9 +273,16 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working. - **Shell** for the times you want a terminal on your phone with no agent at all. It is a genuinely useful mode, not a fallback. +## Pointing one at your own server + +Most of these harnesses can also run against a custom OpenAI-compatible endpoint instead of +their native cloud backend, for one session at a time, an opt-in feature covered in full on +[Custom Model Endpoints](Custom-Model-Endpoints). + ## Read next - [Core Concepts](Core-Concepts) - run modes versus location overlays. +- [Custom Model Endpoints](Custom-Model-Endpoints) - run a harness against your own server. - [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings. - [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode. - [Security](Security) - what skipping permission prompts actually means. diff --git a/docs/wiki/Custom-Model-Endpoints.md b/docs/wiki/Custom-Model-Endpoints.md new file mode 100644 index 00000000..f7e2ed28 --- /dev/null +++ b/docs/wiki/Custom-Model-Endpoints.md @@ -0,0 +1,75 @@ +# Custom Model Endpoints + +Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for +one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM, +a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's +OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering +`GET /v1/models` and `POST /v1/chat/completions` in the standard shape. + +**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**. + +## Adding an endpoint + +Still in App Settings → Models → Custom model endpoints: + +1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`, + say). An API key is optional; most local servers don't check one. +2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it. +3. Pick a **default model** from what was discovered. This is the model the Run-menu entry + below applies with no further choice, so set it once you know which one you want. + +Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker +hosts — these are machine-level infra, not a per-user setting. + +## Running a session against one + +With the setting on and at least one endpoint carrying a usable default model, the **Run** +dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a +custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a +session on that harness exactly the way its own entry would, then points it at the +endpoint's default model. It is a one-off "try this endpoint" action, not a sticky mode — the +plain **Run** button still means "this harness, native cloud" afterward, and a fresh session +never inherits whatever the last one was pointed at. + +Applying a selection **restarts the harness's process in place** — same tab, same +conversation where the harness supports resuming one, fresh environment. That restart is +necessary, not incidental: every supported harness reads its endpoint config at process +start, never per turn, so there is no live hot-swap while a turn is running. + +Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for +redirecting those hasn't landed yet, see below. + +## Which harnesses actually work + +| Harness | Status | +| ------- | ------ | +| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. | +| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. | +| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. | +| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. | +| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. | + +Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry, +not a fixed list here, so this table can go stale before this page does — a greyed-out or +missing entry is the more current answer. + +## What it does not do + +- **No remote or Docker sessions yet.** Both restart their agent differently under the hood + (reattaching a durable tmux session rather than relaunching the process), so redirecting + them needs its own plumbing that hasn't been built. +- **No live hot-swap mid-conversation.** Applying a selection always restarts the process. +- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any, + never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate, + explicit choice per session. + +## Security + +An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save +time and against the address it actually resolves to), the same guard Web Tabs uses for +saved dashboards. Endpoint records and any per-session config files a harness needs are +written with owner-only permissions. See +[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md) +in the repository for the full design reasoning, including why this feature closed a +pre-existing gap in how session environment overrides were guarded rather than opening a new +one. diff --git a/docs/wiki/Settings-Reference.md b/docs/wiki/Settings-Reference.md index 3255f56f..14b04056 100644 --- a/docs/wiki/Settings-Reference.md +++ b/docs/wiki/Settings-Reference.md @@ -92,6 +92,10 @@ Model and effort are both **soft defaults**: the model is written into the case' `.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort` inside a session override them at any time. +**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching +section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server +instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints). + ### Agents & CLIs | Setting | Notes | diff --git a/docs/wiki/_Sidebar.md b/docs/wiki/_Sidebar.md index af84a756..2e6d54b5 100644 --- a/docs/wiki/_Sidebar.md +++ b/docs/wiki/_Sidebar.md @@ -12,6 +12,7 @@ - [The Dashboard](The-Dashboard) - [Agent CLIs](Agent-CLIs) +- [Custom Model Endpoints](Custom-Model-Endpoints) - [Working With Files](Working-With-Files) - [Input And Voice](Input-And-Voice) - [Mobile Guide](Mobile-Guide)