Repository navigation
Relative URLs in WHATWG URL API #12682
Description
Activity
- addedwhatwg-urlIssues and PRs related to the WHATWG URL implementation.Issues and PRs related to the WHATWG URL implementation.
on Apr 27, 2017 I do not suspect that we will be able to actually deprecate
url.parse()for quite some time. For the foreseeable future (at least through 10.x) the two implementations will coexist side-by-side.I vote for "do nothing", at least until it's clear what the use cases for these base-less relative URLs are and how people use them.
If we find out people use them a lot, I think a better solution would be (preferably in user-land) a "RelativeURL" class, not "TolerantURL", which only has pathname/search/searchParams/hash/toString(). Someone would have to specify how this works, but maybe as a first-pass it could have an internal real-URL and parse against
https://example.com/, then just re-expose the pathname/search/searchParams/hash.Reacted by Jessica Franco, Felix Becker, Israeltheminer and codethief@domenic it's worth pointing out that working with relative URLS is extremely common in Node - the most common use case I can think of is when an incoming HTTP request arrives - you get a relative URL under
request.urlwhich you typicallyurl.parse.I think it would be a shame to keep two URL APIs just for relative URLs. It would be really nice if
URLsupported relative URLs somehow - how set in store is the spec at this point regarding that?Reacted by Nathan H. Leung, Jelle De Loecker, Yaroslav Kiliba, Ian Walter, Sergei Startsev, Boris Serdiuk, Zbyszek Tenerowicz, Alex Manning, Dario Vladović, Taka Okunishi and 28 moreIt just doesn't make sense to have a single API for both relative and absolute URLs---the components and parsing rules are far too different depending on the base used for the rest of the API to make any sense. So I'm pretty sure the spec isn't going to change, just because there's no underlying model that makes sense.
The best you can do if you want to use one API is to make up a base URL.
Reacted by Haroen Viaene and Daniel Arthur GallagherReacted by Adrian and Alwin Blok@domenic I realize that this is a hard problem, and one I do not understand very well - but I think having a base URL in Node but not in the browser could cause a lot of incompatibility when people expect code using the same spec to run the same way on both platforms.
I think we can only change
URLto addexample.comor something similar as a default if browsers do. If that isn't practical - then Node should stop callingurl.parsea legacy API and embrace using it for relative URLs.Reacted by Jarrod Davis, yashashav_dk, shitpoet, Adrian, Jay Sandoval and BernardWhich can be done, of course, using the information in an HTTP request (in the case of request.url) so I'm not overly concerned with that particular case.
The key challenge, of course, is that without a base, it's impossible to say for sure which rules to apply to the relative bits. We either must provide a base or we must provide an equivalent context in order to properly handle the URL. Otherwise the parsing will be best guess at best.
Reacted by Domenic DenicolaA compromise might be adding
request.parsedURL()or similar (a function to show that it's expensive) which uses the request info + therequest.urlrelative URL string to return a properly parsedURLinstance with the base constructed from the request info.Which can be done, of course, using the information in an HTTP request (in the case of request.url) so I'm not overly concerned with that particular case.
How? An HTTP server is not aware of its host name, may have several dns host names or none. In fact, I'd argue that the end server should not be concerned with what hostname it's using.
Reacted by Jelle De Loecker, Jarrod Davis, Paul Hawxby, Tv, Elmer Bulthuis, szmcdull, Artur Klesun, Cyerra, shitpoet, Adrian and 3 moreIt can make a best guess using the protocol and host header, both of which may be modified of course, but it provides enough context to provide a base URL when parsing the request URI.
I'll not claim to know which, if any, of the 3 suggested solutions are best, but I just wanted to add a few things to the discussion:
I've seen incoming HTTP requests to Node servers that don't have a
Hostheader in the wild a few times. If it's been striped by a proxy or if the client haven't provided it in the first place I don't know, but you can't rely on this header.The 2nd issue is knowing the protocol. This can be inferred by looking at the
request.socket.encryptedboolean, but this is not an exact science as the TLS might have been terminated in a load balancer.Bottom line is that we can only safely get the path (via
request.url). Getting the rest can be worked around, but not in a nice way.Reacted by shitpoetSo what I'm seeing is that what people really want a partial URL parser for is to parse the origin form of the request target in an HTTP request, which is defined as:
origin-form = absolute-path [ "?" query ]We can introduce a
URLAbsolutePathclass to parse from path state for this specific issue. The semantics are fairly well-defined for this specific case, since we know the scheme is always special (HTTP or HTTPS). Beyond that (including relative, query-only, or fragment-only URLs), however, users are on their own.hmm... I'm hesitant to introduce a new class. This could be approximated in userland fairly easily using something like:
const url = new URL(`https://localhost${absolutePath}?${query}`);
Then look at the bits of
urlthat you care about.@jasnell that doesn't look very usable on user input 😕
A userland module can make it more usable. I'd rather avoid adding a convenience class that is not part of the standard
29 remaining items
@styfle want to move that comment over in nodejs/web-server-frameworks#71? It would be a good starter to the conversation I wanted to have there. And I have comments but don't want to hijack this thread to make them.
Reacted by StevenThe
URLAPI appears to charged blindly down the route of strictness and standards and lost a whole bunch of the utility in the process. I'm not saying necessarily that's a bad thing, but I think it makes the case that the 2 API's should probably remain because they serve 2 different purposes.Say I want to just grab the hash portion of a relative URL
url.parse('foo#bar')works just fine,new URL('foo#bar')does not. I don't care about the hostname, or the protocol, or anything like that, I just want an easy way to flexibly parse URL's, especially if those URL's have been inputted by a user. The new API has lost a lot of utility due to the lack of input flexibility.If I want to strictly parse full URL's, fine,
new URLmakes sense. If I want to perform various flexible utility actions on parsed URLs, useurl.parse. The two are distinct in their function and both should remain or the URL API should be made more flexible. There's little point delegating this to third party modules when there's already code to do this that's being deprecated in favour of a new API that serves and entirely different purpose.Reacted by Gurpreet Atwal, Vitaliy Potapov, Felix Haus, Matthias, szmcdull, Tim, Artur Klesun, shitpoet and Mikael Finstad- added a commit that references this issue
on Mar 3, 2021 Check this one
https://nodejs.org/dist/latest-v14.x/docs/api/all.html#http_message_urlTo parse the URL into its parts:
new URL(request.url, `http://${request.headers.host}`);Once URL object is created this way we can use all its methods and properties.
Ref. Link: https://nodejs.org/dist/latest-v14.x/docs/api/url.html#url_the_whatwg_url_apiGiven that we've (a) added documentation illustrating how to better handle relative URLs with the WHAT-WG API, and (b) We've backed off the deprecation of the legacy API, I'm going to close this issue for now. There's still an argument that could be made on the standards level for more ergonomic handling of relative URLs but those discussions are better directed to the whatwg/url repository.
For the sake of completeness here is the corresponding issue in the whatwg/url repository: whatwg/url#531
Reacted by Audio and CyerraReacted by StevenI think this is very important. The WHATWG API has been designed to standardise existing browser behaviour, not to be the general URL API for platforms such as NodeJS. This thread shows that this causes issues, but it is a problem not as much with NodeJS as with limitations of the standard.
I can predict that this will cause more problems down the road (and not just in Node) as the WHATWG API is becoming more widespread and people will necessarily hack around it to make it meet their needs.
I recently completed my research on the technical part of the problem by releasing this somewhat low level library. My hope is that the community can use it as a basis for building a number of more polished URL APIs that do support relative URLs whilst maintaining compatibility with URLs as defined in the WHATWG standard. I have one attempt at such an API here (but please, come up with alternatives).
The theory behind it is solid and is (still being) written down here. Any help in motivating the WHATWG to take on the issue of relative URLs is welcome.Reacted by strarsis, Roman Frołow, Jonathan Neal, Dmitry Kirilyuk, Paul Hawxby and CyerraReacted by Timothy Gu, Stefan Langeder and Damien GarridoReacted by strarsis- added a commit that references this issue
on Mar 28, 2022 - added a commit that references this issue
on Mar 29, 2022 (b) We've backed off the deprecation of the legacy API
@jasnell has that been formally declared anywhere? It may have been and I've just missed it. Should I PR the typings to remove the deprecation notice?
https://github.com/DefinitelyTyped/DefinitelyTyped/blob/master/types/node/url.d.ts#L63
https://github.com/DefinitelyTyped/DefinitelyTyped/blob/master/types/node/v16/url.d.ts#L63Yep, if you look here https://nodejs.org/dist/latest-v18.x/docs/api/url.html#legacy-url-api, you'll see that the old API is now explicitly marked "Legacy" rather than "Deprecated" as of Node.js 15.13.0
Reacted by Steven, Paul Hawxby, Andrii Oriekhov and Vitaliy PotapovAs of Node.js 19,
url.parse()is once again Deprecated 😮💨Reacted by 0xB0C5Reacted by Bernard- added a commit that references this issue
on Apr 20, 2024 @jasnell, can you point me to the documentation that explains how to handle relative URLs with the WHATWG URL API please? I have read the docs here but it is unclear to me how to handle a relative url in a library in which it is undefined what context such a relative
http.ClientRequest.urlis used in?Node.js docs I used: https://nodejs.org/dist/latest/docs/api/url.html#the-whatwg-url-api
My library in question: https://github.com/sbrl/powahroot/
We are on the track to slowly deprecate the non-standard
url.parse()(#12168 (comment)) in favor of the new WHATWG standard-based URL API. One use case that currently cannot be migrated over fromurl.parse()is the handling of relative URLs.Background
url.parse()accepts incomplete, relative URLs by filling unavailable components of a URL withnull.On the other hand, the
URLconstructor guarantees that all URL objects are fully complete and valid URLs, which means that it throws an exception in case of relative URLs:WHATWG URL API does have the algorithms necessary to parse relative URLs, however, and that is activated if a
baseargument is provided:It is not always the case that a base URL is available, though.
Possible solutions
Do nothing
What this entails is that the currently supported ability to parse relative URLs will die as
url.parse()becomes deprecated.Do not deprecate
url.parse(); otherwise do nothingThis is the most obvious actual solution, but from tickets like #12168, I don't see this as a good idea.
Add a non-standard
TolerantURLclassThis could work if we trick the parser into believing we have a legitimate URL, except there are many conditionals in the URL parser algorithm that provide ad-hoc compatibility fixes with legacy implementations. We would have to make a set of opinionated assumptions about the nature of the URL, such as the URL's scheme.
In addition to parsing, the setters will have awkward semantics. Consider the following:
Something else that's better than what I thought of above...