Files
deepseek-v4-quant/memory/2026-05-04-httpx-read-timeout.md
biondizzle f63eed5cfd Purge INT4 references — expert weights are FP4 (E2M1), not INT4
All docs and scripts updated. Historical memory entries annotated.
2026-05-08 02:33:46 +00:00

42 lines
2.6 KiB
Markdown

# Session: 2026-05-04 18:52:38 UTC
- **Session Key**: agent:main:main
- **Session ID**: a2312560-ab59-4ff6-9fd6-b39daaef3e70
- **Source**: webchat
## Conversation Summary
user: A new session was started via /new or /reset. Run your Session Startup sequence - read the required files before responding to the user. Then greet the user in your configured persona, if one is provided. Be yourself - use your defined voice, mannerisms, and mood. Keep it to 1-3 sentences and ask what they want to do. If the runtime model differs from default_model in the system prompt, mention the default model. Do not mention internal steps, files, tools, or reasoning.
Current time: Monday, May 4th, 2026 - 6:50 PM (UTC) / 2026-05-04 18:50 UTC
assistant: Yo Mike. Vinny here — fresh install, no memories yet, blank slate. What are we building?
user: Sender (untrusted metadata):
```json
{
"label": "openclaw-control-ui",
"id": "openclaw-control-ui"
}
```
[Mon 2026-05-04 18:51 UTC] take a gander at `/home/openclaw/dev/api.vultrinference.com/fragile_proxy.py`
Can you explain the time outs to me? I'm trying to understand if `READ_TIMEOUT` is just waiting for the first chunk or ANY chunk
assistant: OK so here's the deal with the timeouts:
**`READ_TIMEOUT = 300`** — This is passed to `httpx.Timeout(read=300)` and it applies to **ANY chunk**, not just the first one. httpx's `read` timeout fires every time the client is waiting for data from the socket. If there's a 300-second gap between *any* two chunks, it'll raise `httpx.ReadTimeout`.
That said, the code *also* has a second, lower threshold that'll trigger before `READ_TIMEOUT` in most cases:
**`STALE_TIMEOUT = 20`** + **`STALE_CHUNKS = 10`** — This is the app-level stale detection. If the stream produces 10 consecutive chunks with no meaningful content (no content, no reasoning, no tool_calls) AND it's been 20+ seconds since the last meaningful chunk, it marks the stream as stale and bails. This catches the "stream is technically alive but spewing garbage" case that `READ_TIMEOUT` wouldn't catch.
So in practice:
| Scenario | What kills it |
|---|---|
| Stream goes totally silent (no chunks at all) | `READ_TIMEOUT` (300s) |
| Stream keeps sending empty deltas | `STALE_TIMEOUT` (20s) |
| Normal, healthy streaming | Neither — stream completes |
The 300s `READ_TIMEOUT` is really a safety net for "vLLM just stopped sending anything mid-stream." The stale check at 20s is what actually catches most broken streams in practice.
**`RETRY_TIMEOUT = 120`** is separate — it's the total timeout for non-streaming retry requests (where you're waiting for the full response, not chunks).