Repository navigation
h11 fails on multiple targets where other HTTP clients work #95
Description
Activity
Can you file individual bugs for the different issues? "h11 should work" is too vague to figure out actual code changes :-). The problem is to figure out what exactly servers are doing that h11 needs to support.
Multiple content-lengths already has an issue here: #92
What on earth is a "600" response? I've never heard of that.
"Malformed data" means that one of h11's parsing regexps failed. Need more details to figure out which one needs to be loosened and how.
"Receive buffer too long" probably means that the headers were >16384 bytes, which is the default
max_incomplete_event_size: https://h11.readthedocs.io/en/latest/api.html#the-connection-object
This is already configurable, though we could potentially change the default if there's a good reason. The current value is pretty arbitrary; I based it on looking at some HTTP servers and picked something in the same ballpark:Lines 23 to 32 in 68e32db
# If we ever have this much buffered without it making a complete parseable # event, we error out. The only time we really buffer is when reading the # request/reponse line + headers together, so this is effectively the limit on # the size of that. # # Some precedents for defaults: # - node.js: 80 * 1024 # - tomcat: 8 * 1024 # - IIS: 16 * 1024 # - Apache: <8 KiB per line> Probably it would make sense to look at clients too, though. Apparently curl has a hardcoded limit of 102400? https://curl.haxx.se/mail/lib-2019-09/0023.html
The 600 response doesn't actually exist in the standard, it's something that whoever configured the server created. Nonetheless, shouldn't be a reason to reject the response.
These were found out using the httpx library. More examples and discussion here: encode/httpx#767
Here's an issue for one possible cause of the "malformed data" – not sure if it's the one you saw or not. (Or maybe you saw multiple, I dunno)
Whoops, I meant this issue: #97
LinkedIn is an example of where the status code >= 600 comes up in the wild. They are (in)famous for returning 999 status codes. See: https://stackoverflow.com/questions/27231113/999-error-code-on-head-request-to-linkedin or just run
curl -I --url https://www.linkedin.com/company/linkedinHi! Triaging older issues — I think the concrete cases listed here have either shipped or have been split into their own tickets, so this umbrella issue can probably be closed.
Going through each example in the report:
Response status_code should be in range [200, 600), not 600— fixed by PR #140 "Expand the allowed status codes to 999", merged 2021-12-23 (commit96c0a33).h11/_events.py:247-253now allows[200, 1000), so the LinkedIn-style999case mentioned downthread works:>>> import h11 >>> h11.Response(status_code=999, headers=[(b'content-length', b'0')], reason=b'') Response(status_code=999, ...)
multiple Content-Length headers— the maintainer already redirected to Handle multiple Content-Length headers with the same value #92 in the first reply; conflictingContent-Lengthvalues are intentionally rejected per RFC 7230 §3.3.2, and duplicates with the same value are deduplicated inh11/_headers.py:172-184.Receive buffer too long— this is already configurable viaConnection(max_incomplete_event_size=...)documented ath11/_connection.py:69and in the public API docs.malformed data— the maintainer redirected to Update header value validation to match WHAT-WG fetch spec #97 (header-value validation) in the thread; that's the place to track loosening, separately from this umbrella.
If I missed a specific concrete failure that's not covered by one of the above (or the linked tickets), happy to dig further with a reproducer. Otherwise would you mind closing this out?
Disclosure: I drafted this comment with help from Claude Code while triaging stale issues; the PR/commit references and the 999-status reproducer above were verified manually against current
master.
Considering for example this snippet of code
Other errors of the same kind that I've encountered include:
These are all basically ill-configured servers, sometimes even against protocol specs, but they actually appear a lot in the wild. I think these should work nonetheless, as most HTTP clients don't make these kinds of restrictions and they do allow users to see the underlying data despite the misconfigurations.