mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-10 01:09:43 +02:00
More from the review of 5fc391a4, all documentation rather than behaviour.
The changeset was 1602 words of development log, written as the PR grew, with bullets
and loose paragraphs interleaved. That text becomes CHANGELOG.md and the GitHub release
body verbatim, so it is now one user-facing account of what the feature does and what
the real-server work bought, at roughly a fifth the length.
docs/api-reference.md promised a `cmd` field on running-status that the route
deliberately strips (it carries model paths and can carry --api-key).
Two places claimed the apply routes validate `modelId` against the endpoint's
discovered models. Neither does. Dropped the claim rather than adding the check:
discovery can be up to five minutes stale, so a 400 there would refuse a launch that
actually works, and a typo'd id already fails on the CLI's own first request. CLAUDE.md
now says so explicitly, since the absence is the surprising part.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
36 lines
2.1 KiB
Markdown
36 lines
2.1 KiB
Markdown
---
|
|
'aicodeman': minor
|
|
---
|
|
|
|
feat(custom-model): pick a custom endpoint straight from the Run menu
|
|
|
|
#393 landed the backend for custom model endpoints and left it reachable only over the
|
|
HTTP API. This is the rest of it. Turn on Custom model endpoints in App Settings, save
|
|
an endpoint, and the Run dropdown grows a Custom Endpoints section built live off the
|
|
CLI registry, one entry per harness that can actually redirect plus each endpoint you
|
|
saved. Pick one and it launches that harness pointed at your server, asking which model
|
|
first when the endpoint has more than one. Endpoints re-discover themselves every five
|
|
minutes, and one unreachable endpoint never blocks the others. App Settings gains full
|
|
add, edit and delete for endpoints.
|
|
|
|
Seven of the harnesses (opencode, Codex, Gemini, Pi, Grok, DeepSeek and OMP) now launch
|
|
directly onto the endpoint with no restart at all, where before you watched a native
|
|
boot followed immediately by a second one. Claude still launches and then restarts in
|
|
place, which its own resume makes far less jarring.
|
|
|
|
Most of this release's work went into things that only show up against a real server,
|
|
and each was found that way rather than in tests: a freshly launched CLI reporting
|
|
itself busy for its own startup and getting refused; Claude Code assuming a large
|
|
context window for a model it does not recognise and silently overflowing a small one;
|
|
a model whose real context is below what Claude Code's own system prompt costs, which
|
|
no setting can fix and which now warns before launching into a certain failure; and the
|
|
big one, llama.cpp running exactly one model at a time, so applying a selection can
|
|
unload the model another session is using. That last case now asks first, tells you
|
|
which session it affects, and keeps a "loading model" notice on screen for the whole
|
|
swap window, so a prompt sent mid-swap reads as loading rather than as an answer from
|
|
whatever was loaded a moment ago. A background sweep also catches the reverse: your
|
|
session's model being evicted later by somebody else's ordinary use.
|
|
|
|
Remote SSH and Docker sessions are refused for now, since their restart reattaches a
|
|
durable tmux rather than relaunching the agent.
|