mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
docs(wiki): document Custom Model Endpoints and the Run-menu picker
New docs/wiki/Custom-Model-Endpoints.md (auto-synced to the live GitHub wiki on push to master, per docs/wiki/Contributing.md) covers turning the feature on, adding an endpoint, the Run-menu picker's one-off-run behaviour, the per-harness confidence table, and what it deliberately does not do yet (remote/Docker sessions, live hot-swap). Linked from the sidebar, from Agent-CLIs.md's "Read next" list plus a short pointer section, and from Settings-Reference.md's Models section. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RqZeHrRS6DYcGcGX2p9EwG
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
25fae9ad10
commit
98d26e14d9
@@ -273,9 +273,16 @@ into the case's `.claude/settings.local.json` so that `/model` keeps working.
|
||||
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
|
||||
genuinely useful mode, not a fallback.
|
||||
|
||||
## Pointing one at your own server
|
||||
|
||||
Most of these harnesses can also run against a custom OpenAI-compatible endpoint instead of
|
||||
their native cloud backend, for one session at a time, an opt-in feature covered in full on
|
||||
[Custom Model Endpoints](Custom-Model-Endpoints).
|
||||
|
||||
## Read next
|
||||
|
||||
- [Core Concepts](Core-Concepts) - run modes versus location overlays.
|
||||
- [Custom Model Endpoints](Custom-Model-Endpoints) - run a harness against your own server.
|
||||
- [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings.
|
||||
- [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode.
|
||||
- [Security](Security) - what skipping permission prompts actually means.
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
# Custom Model Endpoints
|
||||
|
||||
Point a harness at your own OpenAI-compatible server instead of its native cloud backend, for
|
||||
one session at a time. "Custom endpoint" covers **local** hardware (llama.cpp, Ollama, vLLM,
|
||||
a home GPU rig, DGX Spark, Strix Halo) and **cloud** services (Azure AI Foundry's
|
||||
OpenAI-compatible endpoint, OpenRouter, a company gateway) alike, anything answering
|
||||
`GET /v1/models` and `POST /v1/chat/completions` in the standard shape.
|
||||
|
||||
**Off by default.** Turn it on in App Settings → Models → **Custom model endpoints**.
|
||||
|
||||
## Adding an endpoint
|
||||
|
||||
Still in App Settings → Models → Custom model endpoints:
|
||||
|
||||
1. **+ Add endpoint** — give it an id, a label, and the base URL (`http://192.168.1.50:8080`,
|
||||
say). An API key is optional; most local servers don't check one.
|
||||
2. **Discover** — fetches the endpoint's own model list over `GET /v1/models` and stores it.
|
||||
3. Pick a **default model** from what was discovered. This is the model the Run-menu entry
|
||||
below applies with no further choice, so set it once you know which one you want.
|
||||
|
||||
Endpoint management is admin-only in multi-user mode, the same as remote hosts and Docker
|
||||
hosts — these are machine-level infra, not a per-user setting.
|
||||
|
||||
## Running a session against one
|
||||
|
||||
With the setting on and at least one endpoint carrying a usable default model, the **Run**
|
||||
dropdown grows a **Custom Endpoints** section: one entry per harness that can redirect to a
|
||||
custom endpoint, per saved endpoint, e.g. "Claude Code (llama.cpp)". Picking one starts a
|
||||
session on that harness exactly the way its own entry would, then points it at the
|
||||
endpoint's default model. It is a one-off "try this endpoint" action, not a sticky mode — the
|
||||
plain **Run** button still means "this harness, native cloud" afterward, and a fresh session
|
||||
never inherits whatever the last one was pointed at.
|
||||
|
||||
Applying a selection **restarts the harness's process in place** — same tab, same
|
||||
conversation where the harness supports resuming one, fresh environment. That restart is
|
||||
necessary, not incidental: every supported harness reads its endpoint config at process
|
||||
start, never per turn, so there is no live hot-swap while a turn is running.
|
||||
|
||||
Entries are hidden entirely for a session in a **remote (SSH) or Docker case** — support for
|
||||
redirecting those hasn't landed yet, see below.
|
||||
|
||||
## Which harnesses actually work
|
||||
|
||||
| Harness | Status |
|
||||
| ------- | ------ |
|
||||
| **Claude Code, opencode, Pi, Grok, OMP** | Verified end-to-end against a real local server. |
|
||||
| **Codex** | Config is correct, but Codex only speaks the Responses API, which llama.cpp-style servers don't implement. A protocol gap, not a Codeman bug. |
|
||||
| **Gemini** | Fails with an auth error gemini-cli raises once redirected. Unresolved; don't rely on it yet. |
|
||||
| **DeepSeek** | Reaches the server but gets a consistent 404. Root cause not identified. |
|
||||
| **Antigravity** | No known custom-endpoint mechanism at all. Not offered. |
|
||||
|
||||
Which harnesses show up in the Run-menu picker is read live off Codeman's own CLI registry,
|
||||
not a fixed list here, so this table can go stale before this page does — a greyed-out or
|
||||
missing entry is the more current answer.
|
||||
|
||||
## What it does not do
|
||||
|
||||
- **No remote or Docker sessions yet.** Both restart their agent differently under the hood
|
||||
(reattaching a durable tmux session rather than relaunching the process), so redirecting
|
||||
them needs its own plumbing that hasn't been built.
|
||||
- **No live hot-swap mid-conversation.** Applying a selection always restarts the process.
|
||||
- **Nothing is shared with your real cloud credentials.** The endpoint's own key, if any,
|
||||
never touches your Anthropic/OpenAI/Google login — a custom endpoint is a separate,
|
||||
explicit choice per session.
|
||||
|
||||
## Security
|
||||
|
||||
An endpoint's base URL can't point at a link-local or cloud-metadata address (both at save
|
||||
time and against the address it actually resolves to), the same guard Web Tabs uses for
|
||||
saved dashboards. Endpoint records and any per-session config files a harness needs are
|
||||
written with owner-only permissions. See
|
||||
[custom-model-endpoints-plan.md](https://github.com/Ark0N/Codeman/blob/master/docs/custom-model-endpoints-plan.md)
|
||||
in the repository for the full design reasoning, including why this feature closed a
|
||||
pre-existing gap in how session environment overrides were guarded rather than opening a new
|
||||
one.
|
||||
@@ -92,6 +92,10 @@ Model and effort are both **soft defaults**: the model is written into the case'
|
||||
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
|
||||
inside a session override them at any time.
|
||||
|
||||
**Custom model endpoints** (off by default) adds a saved-endpoint list plus a matching
|
||||
section to the Run dropdown, for pointing a harness at your own OpenAI-compatible server
|
||||
instead of its native cloud backend. See [Custom Model Endpoints](Custom-Model-Endpoints).
|
||||
|
||||
### Agents & CLIs
|
||||
|
||||
| Setting | Notes |
|
||||
|
||||
@@ -12,6 +12,7 @@
|
||||
|
||||
- [The Dashboard](The-Dashboard)
|
||||
- [Agent CLIs](Agent-CLIs)
|
||||
- [Custom Model Endpoints](Custom-Model-Endpoints)
|
||||
- [Working With Files](Working-With-Files)
|
||||
- [Input And Voice](Input-And-Voice)
|
||||
- [Mobile Guide](Mobile-Guide)
|
||||
|
||||
Reference in New Issue
Block a user