Hacker Newsnew | past | comments | ask | show | jobs | submit | phiresky's commentslogin

The EU told them they can not by default prefer their own maps product when linking from their search product, they have to allow the user to choose.

They could have simply added a selector when a user first clicks on the maps preview in the search result, and then remembered it on device or across that user's account.

But of course, then the user could choose a competitor's product and Google would have to honor it. That would be horrible, so instead they just made the user experience worse for everyone and elegantly made people blame the EU.


I wonder if the EU will someday figure out a patch for this "Deteriorate everyone's user experience in a way that is technically compliant and blame the EU" behavioural exploit...

I'm blaming Google for this, but most people just don't know why it happens. Probably they'll end up mandating that Google must show a notice with the reason and it will be implemented in the worst possible way.

All of these hissy fits Google and Apple are throwing are just making me wanting to use them less and less and I'm doing my best to explain how tech companies are doing malicious compliance to others.


How about EU starts legistlating towards a competitive market and creates its own web search Tech Giant ?

So Google is forced to take competitor to account or will be pushed out of the market ?

No? Better to complain and do nothing positive? Welcome to EU Parliment.


The European Commission is way too pro-business for that, even when the said businesses are parasitic actors who don't even pay taxes in the EU.

Plus that just isn't how governments work. Governments work by imposing costs on others. They do not build.

What does it has to do with the topic? And saying “governments do not build” is a crazy antihistorical take: many things, from the space industry to the telco one were built by governments, as well as the energy and transport sectors in many countries.

For the past 50-ish years the West convinced itself that government should not build, and as a result we're suffering from a slow collapse of basic infrastructures. Meanwhile, in China and the Gulf, government kept building and the world's center of gravity steadily keeps moving eastwards.


"What have the Romans ever done for us?"

I merely illustrate that the EU commission certainly has the means to replace Google in the EU. Build a new Youtube. Build the services that society, frankly, depends on now ... which is quite a few of them.

And they're not doing that. They are only making it harder for existing services to exist, in order to change the situation, and in order to make money for themselves. But they're not helping in the sense of building anything.


Yeah, it's very overpowered and should be nerfed - they still haven't patched the cookie consent dialogs out.

To expand upon the silly comment a bit, not sure about this particular initiative, but to me it very much feels like something that should be configurable at browser level, instead of half-assed dark-pattern-filled JS dialogs: https://killthecookiebanner.eu/


It definitely should be, but this just goes to show that regulations are necessary if we don't want companies to unnecessarily screw with users. It would have been better that way for every user, but it would be worse for people who want to sell everyone's personal information.

Oh, and please ignore that this surveillance apparatus is now being used to expell people, all of which already has resulted in many missing people, and even some deaths. Any historical parallels are entirely coincidental.


It's fine + obligation to fix, and daily fine until fixed, and higher fine if you do it again.

I know that sounds like it comes from "it will never work", but that's how active directory, office file format, etc ... got opened, and in a case more similar to here that's how windows n and decoupling of media and internet component from windows internals happened.

The issue is, yes, it takes time.

Another issue, purely PR, is that yes the EU gets attacked repeatedly by people trying to match "respect the law of the market or leave" with "the EU tax US companies because they can't compete".


> It's fine + obligation to fix, and daily fine until fixed, and higher fine if you do it again.

The important proviso is that the penalties need to (on some non-geological timescale) get high enough that the offender is no longer capable of operating. That could mean the fines reach $500 billion, or it could mean the company is barred from operating in the EU, or it could mean people get arrested and assets are seized. But unless the penalties become crippling, it won't matter. It needs to reach a point where the downsides of noncompliance are actually greater than the benefits.


NIS2 allows for the arrest of managers in case of cybersecurity incidents resulting from negligence. I think it's a step in the right direction, but we'll have to see how it plays out.

In the EU they like the "daily fine until fixed". It's how Microsoft bowed down.

But unlike the story some on HN and in the US like to think, US tech companies don't ever get there except rare exception, they know the game.


The solution is to mandate that websites are built for browsers/devices as much as they are for humans. That way you never see a cookie banner because you've set up your device to never accept marketing cookies; A Google Maps link clicks through to Google Maps because you told your browser Google Maps was your preferred online mapping tool.

The corrective force should be market competition, but Google is a monopoly not a monopoly with so many moats it might as well be a monopoly.

They also buy up the competition before it gets a chance as is the way now.


Why should be? I think it is abundantly clear by now that market competition alone will not fix this. Though that does not stop people from claiming "Yes that failed and people suffered/died, but that was not real capitalism! If only we had a real free market, it would have been better!"

The specific problem with Maps is that Google _can't_ deep link to other providers except by lat-lng or address, both of which throw away the actual context of the PoI in the common case of searching for PoIs. Well, it could perhaps deep-link to OSM-based maps, since their map data is open... but Google, as a proprietary map maker, probably treats OSM data as radioactive.

They could make available a URI using the geo: scheme, but the hacky ways of including a POI as part of a geo: URI (ie using a query string) feel generally unsatisfactory.

Google measures latency at microseconds.100 ms ≈ 0.2% fewer searches. Google tracks this very closely and invests a lot as that's the operational life blood of the company. You can make more money by simply making search easier.

To save "it's easy" and Google chooses to nuke their own products for... I don't know what, stick it the European regulators, strikes me as incorrect. It's because it's incredibly hard to navigate this arbitrary punative regulatory framework.

No one would use the other maps because other maps are worse. As much as euro regulators hem and haw about lack of competition, it's precisely these types of dictats that make Europe not have any viable competitors in tech.


Google also makes money by making search worse.

https://www.wheresyoured.at/the-men-who-killed-google/


[flagged]


I often hate how loaded/incendiary Ed Zitron's tone is and don't like some of his takes, but I wouldn't call it "some random blog". I run a random blog, that guy's closer to being a proper journalist: https://en.wikipedia.org/wiki/Ed_Zitron

Referencing a blog is fine. But when you're on a forum discussing ideas just forwarding a blog link is bad taste. Summarize the key arguments. Do some work.

It's like if someone is arguing free markets and I just provide a link to a WSJ editorial. It's rude, requires no work on the sender and imposes work on the receiver to read and synthesize and respond to the argument. I see this occasionally where someone will just have an LLM generate a response (which it will happily do for any topic taking any position) and just passing it off.


> Summarize the key arguments. Do some work.

I'd still take at least a blog post with relevant information over nothing, but yeah, that is fair.


> No one would use the other maps because other maps are worse.

Do you realize how many commercial driving apps are based around OSM data?


It's malicious compliance with the intent of having the users misunderstand what's happening (that you don't have to do that but do it on purpose, and instead blame the law protecting them).

And eh, it works for the cookie banner so why change the strategy ? The comment you're answering to prove it still works. As long as people can't be bothered to think and inform themselves for a second about it, they won't stop. Same with the purposefully annoying and unclear and user hostile 'gdpr popup' about third party consent.

I'm not saying I agree with it. I'm saying from Google POV that's the correct move given their intent, the issue then become people like parent.


Yes, but by the time you are sued and fined into compliance again, you already made a ton of money to pay 10x the next fine.

Yes and no, it's true now but the EU fines are not just "you did that" but "you did that we're going to assume you're going to be better and if you do again next time it will really hurt". Which I think is a fair point of view.

You might still answer too little to late, but given that we're at the starting point to a massive sovereignty push here thanks to the current US admin being proved as a possibility rather than a novelty, I think this matters a lot.


Well no, the cookie banner is a well intentioned but flawed and poorly designed law. Any law that is that technically specific but relies on people without an understanding of how it works is doomed to repeat the same fate. The eu should have seen prop 65 and not sleepwalked back into it.

1. The EU works on intent of the law rather than specific wording and precedence.

2. The law is a very simple "if you want to do it, you have to make sure they know", intended to force information without creating excessive administrative / legal / tech burden, please inform me about what other way you would present it that would reach the goal without being subject to malicious compliance ?


I understand. I'm generally a proponent of the spirit of the law rather than the wording of the law, and I've argued that is how the EU works here in the past. I think GDPR is a much, much better attempt at solving the problem, and has made meaningful change in tech industries. The ePrivacy directive just added pseudo-mandatory popups to every company without a technical lawyer.

Why? The whole point of the GPDR was to prevent medical information being used for insurance and all sorts of purposes.

Then come the lists of what exceptions are approved. Your medical info is used for divorces (anything involving court cases, anything involving criminal law), the police has access to it, your mayor has access to it, tax departments have access to it (think you can not pay tax and pay for your kid's cancer treatment instead? In Europe, think again). Insurance (if you get treatment for getting hurt in traffic your car insurance goes up). Unemployment (if you get treated for anything drug-related ...). Hospitals and doctors can use your medical information without your permission (for billing, for other treatments, for deciding if you should be interned, ...). And so on and so forth.

Oh and there are even silent exceptions. You see, YOU can't sue anyone under the GPDR. You can only ask a specific "supervisory authority" (you can't even choose which one)

They are under control of the executive, and so it is in most cases the currently elected party that decides if your GPDR complaint does anything, NOT the courts. Not the police. Not the public prosecutor. None of that. And it's even closed on the back end: you don't agree with these "supervisory authority"'s actions? Doesn't matter if you're complainant or defendant. You can't sue them either. You can't get a judge on your case, only appointed politicians.

There are even organizations that the GPDR supposedly applies to that have their own supervisory authority. Interpol violated your rights? No worries, file your complaint here in this building. You know, the building with "Interpol" on it in big letters.

So really, we do not even know the full list of exceptions.

More generally, the GPDR was supposed to prevent further encroachment of all sorts of organizations on privacy, with a big focus on medical data. It has achieved the opposite of that. FOR NOW (and not in every country) the only way to get a private medical file is to only use private medical care. For now that is still possible.

It's like the DMA (Digital Markets Act). Prevents organizations from using control of the OS to implement policy. There's a few exceptions though. Google gets an exception. Apple gets an exception. Through specific deals made with these organizations and the EU commissioner of

Nobody seems to have thought to scream into the commissions face: "THEN WHAT'S THE POINT?".

Well, who made those deals? Thierry Breton. He currently serves as a remunerated member of Bank of America’s Global Advisory Council (who have huge investments in Alphabet and Apple).

Yeah, I get why you want to focus on the intent only and not on what practically happened. Theory and practice are very, very different and the EU is incredibly pro-business and uses their power to literally grant billionaires exceptions to laws. That's how Goldman Sachs got it's first communist president (Barosso, who saved Goldman Sachs as president of the EU commission). That's reality, but of course the intent is thoroughly disguised, and you don't want to talk about the difference.


Gdpr whole point was about insurance? Can you please diversify your news source and educate yourself? I didn't even bother to read your pamphlet of a comment after seeing such an obviously wrong first sentence

No it was about privacy. Specifically given as an example in the actual law, privacy of medical data FROM insurance companies. But of course privacy from everyone.

The goal was not even remotely achieved, and this was 100% on purpose.


There a reason why we have courts. Laws aren’t algorithms despite what the tech world wet dream might want them to be.

Then again, the very same people complain about GDPR not being technically super specific.

At some point we have to accept the pattern. People with ideological objections against any limits at all will try to frame any regulation as stupid, regardless of what is in it.


They could also just have their maps and others show in the results as normal.

That would suck, though? Showing places on a map makes so much more sense than burying them in a list of webpages.

Google really sucks at giving users a choice. The Gmail app on iOS has a browser selector dialog, which asks if you want to open a link in an email in Chrome or Safari. Despite the fact that you can change your default browser in iOS settings if you really want to use Chrome. And even if you choose Safari and don’t ask it to remind you every time, it will still prompt you again later if you’d like to switch to Chrome this time. Couldn’t I just pick Safari and you respect that choice until I say otherwise?

I do wonder if people using Chrome on iOS face the same issue, or if Google respects your browser of choice a bit more if it’s their browser.


~don't~ be evil

Probably every Google team has an enshittification expert.

I'm a bit disappointed that this only solves the "find index of file in tar" problem, but not at all the "partially read a tar.gz" file problem. So really you're still reading the whole file into memory, so why not just extract the files properly while you are doing that? Takes the same amount of time (O(n)) and less memory.

The gzip-random-access problem one is a lot more difficult because the gzip has internal state. But in any case, solutions exist! Apparently the internal state is only 32kB, so if you save this at 1MB offsets, you can reduce the amount of data you need to decompress for one file access to a constant. https://github.com/mxmlnkn/ratarmount does this, apparently using https://github.com/pauldmccarthy/indexed_gzip internally. zlib even has an example of this method in its own source tree: https://github.com/gcc-mirror/gcc/blob/master/zlib/examples/...

All depends on the use case of course. Seems like the author here has a pretty specific one - though I still don't see what the advantage of this is vs extracting in JS and adding all files individually to memfs. "Without any copying" doesn't really make sense because the only difference is copying ONE 1MB tar blob into a Uint8Array vs 1000 1kB file blobs

One very valid constraint the author makes is not being able to touch the source file. If you can do that, there's of course a thousand better solutions to all this - like using zip, which compresses each file individually and always has a central index at the end.



This is very cool. Worth a submission by itself.


> Apparently the internal state is only 32kB

Exactly. And often this state is either highly compressible or non-compressible but only sparsely used. The latter can then be made compressible by replacing the unused bytes with zeros.

Ratarmount uses indexed_gzip, and when parallelization makes sense, it also uses rapidgzip. Rapidgzip implements the sparsity analysis to increase compressibility and then simply uses the gztool index format, i.e., compresses each 32 KiB using gzip itself, with unused bytes replaced with zeros where possible.

indexed_gzip, gztool, and rapidgzip all support seeking in gzip streams, but all have some trade-offs, e.g., rapidgzip is parallelized but will have much higher memory usage because of that than indexed_gzip or gztool. It might be possible to compile either of these to WebAssembly if there is demand.


> Each seek point is accompanied by a chunk (32KB) of uncompressed data which is used to initialise the decompression algorithm, allowing us to start reading from any seek point.

> Apparently the internal state is only 32kB, so if you save this at 1MB offsets, you can reduce the amount of data you need to decompress for one file access to a constant.

You may need to revisit the definition of a constant. A 1/32 additional data is small but it still grows the more data you’re trying to process (we call that O(n) not O(1)). Specifically it’s 3% and so you generally want to target 1% for this kind of stuff (once every 3mib)

And the process still has to read though the enter gzip once to build that index


I think you're looking at a different perspective than me. At _build time_ you need to process O(n), yes, and generate O(n) additional data. But I said "The amount of data you need to decompress is a constant". At _read time_, you need to do exactly three steps:

1. Load the file index - this one scales with the number of files unless you do something else smart and get it down to O(log(n)). This gives you an offset into the file. *That same offset/32 is an offset into your gzip index.*

2. take that offset, load 32kB into memory (constant - does not change by number of files, total size, or anything else apart from the actual file you are looking at)

3. decompress a 1MB chunk (or more if necessary)

So yes, it's a constant.


My bad. Yes from decompression perspective you have O(1) ancillary data to initiate decompression at 1 seek point.

This is how seeking can work in encrypted data btw without the ancillary data - you just increment the IV every N bytes so there’s a guaranteed mapping for how to derive the IV for a block so you’re bounded by how much extra you need to encrypt/decrypt to do a random byte range access of the plaintext.

But none of this is unique to gzip. You can always do this for any compression algorithm provided you can snapshot the state - the state at a seek point is always going to be fairly small.


I actually first thought this wasn't possible at all because I'm used to zstd which by default uses a 128MB window and I usually set it to the max (2GB window). 32kB is _really_ tiny in comparison. On the other hand though, zstd also compresses in parallel by default and has tools built in to handle these things, so seekable zstd archives are fairly common.


If anyone knows a similar solution for zstd, I'm very interested. I'm doing streaming uncompression to disk and I'd like to be able to do resumable downloads without _also_ storing the compressed file.


https://github.com/martinellimarco/indexed_zstd

https://github.com/martinellimarco/libzstd-seek

Note, however, that this can only seek to frames, and zstd still only creates files containing a single frame by default. pzstd did create multi-frame files, but it is not being developed anymore. Other alternatives for creating seekable zstd files are: zeekstd, t2sz, and zstd-seekable-format-go.


Thanks, this is helpful. I might just end up using content defined chunking in addition/instead, but it's good to know that there is a path forward if I stick with the current architecture.


Tar doesn't need to imply gzip (or bzip2, or zstd, etc). Tar's default operation produces uncompressed archives.


There's a bit of an issue with the linked deployment (in my opinion). In the most zoomed out view you should see the first layer of blocks - very big blocks titled "English language", "French language", "German language". See https://phiresky.github.io/isbn-visualization/ maybe. That makes it a bit easier to read.

The point of the visualization is showing different attributes of books in the space of ISBNs. ISBNs correlate with country, publisher, and release date, that's why using it as a space is useful. You can clearly see the history of when blocks were created, which blocks are rarer than others (present in fewer libraries), and (on the AA hosting) which blocks are more present in AA vs not.

In any case though, yes ISBNs as spatial data are clearly not perfect. Do you have any suggestions that would order the 100 million data points better?


Here's my article on how I built it - and also an instance hosted on GitHub pages if the AA domain is blocked for you: https://phiresky.github.io/blog/2025/visualizing-all-books-i...

Happy to answer questions as always :)


Love this!


Netdata used to be really impressively minimal, performant, and packed with functions. Fully GPL open source. You ran one install command and it started a web-ui at localhost:19999 in a few seconds. The UI loaded instantly and had hundreds of graphs. You could tell the author was a single opinionated person obsessed with the maximum of monitoring with the minimum footprint.

It auto-detects many programs like docker, nginx and postgresql and automatically creates dashboards for them. It also has many dashboards about system internals I didn't even know were great to monitor, so it taught me a lot. For example, seeing a CPU pinned at 100% processing interrupts because of a network interface overload or having time frames with high IOwait during a SQL query clearly meaning there's some larger seq scans happening.

You also needed zero configuration, no login, etc.

Then they added multi-instance monitoring purely client side - the browser remembers other instance domains and links between them - pretty neat and completely uninvasive.

Then they introduced their cloud login, where you can monitor multiple instances remotely/together. They had a `--no-cloud` flag though if you did not want it. But by now they've removed that flag and they say patching out the cloud functionality is bypassing their license [1]. Some functionality is locked behind premium upgrades, and you get prevented from adding more than N metrics or M instances. It's still _possible_ to use netdata without going through their cloud but you have to go through a nag window every time you try to open the local UI. It's clear they don't want you to use it anymore, and I don't really feel comfortable about their default auto-updating local install any more either.

Now it's still impressive and useful, but it's much more an enterprise focusued tool than an "i have this server i want to monitor" tool.

Of course I understand they need to make money, but what used to be trivial to understand (hooks into everything in your system it can and opens a single port to display it) has become a whole huge integrated ecosystem and for me personally it's competing in the space where I'd probably rather spend the time to make a proper Prometheus/Grafana setup instead.

[1] https://github.com/netdata/netdata/discussions/17594#discuss...


Totally agree. I was quite impressed with UI and bling bling when one of our employees installed it to monitor one of our servers (university team of 15 people) .However it turned out to be a total cognitive overload and very hard to bring it down to the strictly necessary to maintain the server. Then the cloud login stuff was a show stopper. We moved from nagios via icinga to checkmk, but we are still quite unhappy with good metrics based monitoring (we had munin at some point). A lot of the solutions seems overkill or oversimplify alarm states leading to a lot of false positives or duplicate notifications.


I used to love netdata, but right around dashboard v3 I switched to Prometheus+grafana

I still miss the no-fuss configuration and the anomaly detection was well done, but I just can't do the required cloud thing it does now


A $120M spend on AWS is equivalent to around a $12M spend on Hetzner Dedicated (likely even less, the factor is 10-20x in my experience), so that would be 3% of their revenue from a single customer.


> A $120M spend on AWS is equivalent to around a $12M spend on Hetzner Dedicated (likely even less, the factor is 10-20x in my experience), so that would be 3% of their revenue from a single customer.

I'm not convinced.

I assume someone at Netflix has thought about this, because if that were true and as simple as you say, Netflix would simply just buy Hetzner.

I think there lots of reasons you could have this experience, and it still wouldn't be Netflix's experience.

For one, big applications tend to get discounts. A decade ago when I (the company I was working for) was paying Amazon a mere $0,2M a month and getting much better prices from my account manager than were posted on the website.

There are other reasons (mostly from my own experiences pricing/costing big applications, but also due to some exotic/unusual Amazon features I'm sure Netflix depends on) but this is probably big enough: Volume gets discounts, and at Netflix-size I would expect spectacular discounts.

I do not think we can estimate the factor better than 1.5-2x without a really good example/case-study of a company someplace in-between: How big are the companies you're thinking about? If they're not spending at least $5m a month I doubt the figures would be indicative of the kind of savings Netflix could expect.


We run our own infrastructure, sometimes with our own fincing (4), sometimes external (3). The cost is in tens of millions per year.

When I used to compare to aws, only egress at list price costs as much as my whole infra hosting. All of it.

I would be very interested to understand why netflix does not go 3/4 route. I would speculate that they get more return from putting money in optimising costs for creating original content, rather than cloud bill.


> I would be very interested to understand why netflix does not go 3/4 route. I would speculate that they get more return from putting money in optimising costs for creating original content, rather than cloud bill.

I invest in Netflix, which means I'm giving them some fast cash to grow that business.

I'm not giving them cash so that they can have cash.

If they share a business plan that involves them having cash to do X, I wonder why they aren't just taking my cash to do X.

They know this. That's why on the investors calls they don't talk about "optimising costs" unless they're in trouble.

I understand self-hosting and self-building saves money in the long-long term, and so I do this in my own business, but I'm also not a public company constantly raising money.

> When I used to compare to aws, only egress at list price costs as much as my whole infra hosting. All of it.

I'm a mere 0,1% of your spend, and I get discounts.

You would not be paying "list price".

Netflix definitely would not be.


Of course netflix is optimising costs, otherwise it would not be a business, I just think they put much more effort elsewhere. They could be using other words, like "financial discipline" :)

My point is that even if I get 20 times discount on egress its still nowhere close, since i have to buy everything else - compute, storage are more expensive, and even with 5-10x discounts from list price its not worth it.

(Our cloud bills are in the millions as well, I am familiar with what discounts we can get)


Even then you can just err.downcast_ref::<std::Io::Error>() though to get the underlying IOError, no?


I'm happy to answer any questions! Nice to see this here again :)


This is great! Especially the DB sync part, because that happens before a user interaction, so you actually have to wait for it (the update itself can run in the background).

It always felt like such a waste to me how the DB always downloads tens of megabytes of data when likely only 1kB has changed. I mean I also really appreciate the beauty of how simple it is. But I'd bet even a delta against a monthly baseline file would reduce the data by >90%.

Also, it would be interesting to see how zstd --patch-from compares to the used delta library. That is very fast (as fast as normal zstd) and the code is already there within pacman.

For the recompression issue, there is some hard to find libraries that can do byte-exact reproducible decompression https://github.com/microsoft/preflate-rs but I don't know of any that work for zstd.


There's an extension to ISO8601 that fixes this and is starting to become supported in libraries:

    2019-12-23T12:00:00-02:00[America/Sao_Paulo]
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: