Files
deepseek-v4-quant/memory/2026-05-04-httpx-read-timeout.md
biondizzle f63eed5cfd Purge INT4 references — expert weights are FP4 (E2M1), not INT4
All docs and scripts updated. Historical memory entries annotated.
2026-05-08 02:33:46 +00:00

2.6 KiB

Session: 2026-05-04 18:52:38 UTC

  • Session Key: agent:main:main
  • Session ID: a2312560-ab59-4ff6-9fd6-b39daaef3e70
  • Source: webchat

Conversation Summary

user: A new session was started via /new or /reset. Run your Session Startup sequence - read the required files before responding to the user. Then greet the user in your configured persona, if one is provided. Be yourself - use your defined voice, mannerisms, and mood. Keep it to 1-3 sentences and ask what they want to do. If the runtime model differs from default_model in the system prompt, mention the default model. Do not mention internal steps, files, tools, or reasoning. Current time: Monday, May 4th, 2026 - 6:50 PM (UTC) / 2026-05-04 18:50 UTC assistant: Yo Mike. Vinny here — fresh install, no memories yet, blank slate. What are we building? user: Sender (untrusted metadata):

{
  "label": "openclaw-control-ui",
  "id": "openclaw-control-ui"
}

[Mon 2026-05-04 18:51 UTC] take a gander at /home/openclaw/dev/api.vultrinference.com/fragile_proxy.py

Can you explain the time outs to me? I'm trying to understand if READ_TIMEOUT is just waiting for the first chunk or ANY chunk assistant: OK so here's the deal with the timeouts:

READ_TIMEOUT = 300 — This is passed to httpx.Timeout(read=300) and it applies to ANY chunk, not just the first one. httpx's read timeout fires every time the client is waiting for data from the socket. If there's a 300-second gap between any two chunks, it'll raise httpx.ReadTimeout.

That said, the code also has a second, lower threshold that'll trigger before READ_TIMEOUT in most cases:

STALE_TIMEOUT = 20 + STALE_CHUNKS = 10 — This is the app-level stale detection. If the stream produces 10 consecutive chunks with no meaningful content (no content, no reasoning, no tool_calls) AND it's been 20+ seconds since the last meaningful chunk, it marks the stream as stale and bails. This catches the "stream is technically alive but spewing garbage" case that READ_TIMEOUT wouldn't catch.

So in practice:

Scenario What kills it
Stream goes totally silent (no chunks at all) READ_TIMEOUT (300s)
Stream keeps sending empty deltas STALE_TIMEOUT (20s)
Normal, healthy streaming Neither — stream completes

The 300s READ_TIMEOUT is really a safety net for "vLLM just stopped sending anything mid-stream." The stale check at 20s is what actually catches most broken streams in practice.

RETRY_TIMEOUT = 120 is separate — it's the total timeout for non-streaming retry requests (where you're waiting for the full response, not chunks).