Initial implementation of claude-max-bridge

OpenAI-compatible FastAPI server wrapping the Claude Code CLI.
Translates /v1/chat/completions requests into claude --print subprocess
calls using stream-json input/output for multi-turn support and real
streaming deltas.

Features:
- Full multi-turn conversation support via --input-format stream-json
- Real-time streaming with --include-partial-messages delta tracking
- Rate limit (429) and auth error (401) detection from CLI stderr
- Configurable timeout, graceful subprocess cleanup on disconnect
- /v1/models endpoint and /health check
This commit is contained in:
2026-05-22 16:08:44 +02:00
parent 9862a9d585
commit 3a58fba521
5 changed files with 599 additions and 1 deletions
+13
View File
@@ -0,0 +1,13 @@
FROM python:3.12-slim
RUN pip install --no-cache-dir fastapi "uvicorn[standard]"
COPY server.py /app/server.py
WORKDIR /app
# claude binary and ~/.claude config dir are bind-mounted at runtime:
# -v /usr/local/bin/claude:/usr/local/bin/claude:ro
# -v /root/.claude:/root/.claude:ro
EXPOSE 8000
CMD ["uvicorn", "server:app", "--host", "0.0.0.0", "--port", "8000"]