Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is not the 90s - monitoring and alerting are far more robust, and supply chains far more just in time.

There’s no reason a modern fleet couldn’t detect failures in hardware and automatically order replacement from Amazon or whatever. This is not to say you won’t need personnel, but you would be surprised how much of the fleet management at AWS and GCP at least are brute forced (as opposed to being completely automatic). To their defense they have a far more complicated and diverse fleet than what I’m describing, which is sort of my point.

99% of customers just need a highly available setup for their stateful boxes (DBs) and some for their stateless services. And the great thing about this setup is you can extend the stateless bit (with the cost of latency) to the cloud more or less infinitely.

As far as the rest of your point about cooling and network - it's a good point, but for a workload that isn't a datacenter, in my experience the majority of outages aren't due to that - it's due to misconfiguration.

Take a personal house. How often does someone break into the average person's home? Pretty unlikely in a low crime area. Take internet. There are some areas in the United States where there are multiple 1Gbs or even 10Gbs internet providers, you could redundantly network. In any case it's not that the cloud shouldn't be used, it's just interesting the direction things are going.

Of course, if money isn't an issue by all means people should use serverless and cloud spanner and be done with it.



> This is not to say you won’t need personnel

Ding ding ding, we have a winner!

Paying premiums for BigCloud means not paying salaries, not worrying about an outage that could happen while your seniors are on vacation, not needing to set up resilient internal processes and controls for datacenter management, not worrying about datacenter management compliance, etc. etc.

Needing to hire datacenter personnel is a problem and it is a problem that BigCloud solved.


You're thinking about it the wrong way. Sure you're not paying salaries, but now you're paying for experts who understand the cloud and you're paying more for those people compared to datacenter people. At Azure I was surprised how many data center people didn't even have a college degree and how little they were paid. Ultimately the only thing that matters is the total cost. To take true advantage of the cloud you need to use cloud specific technologies - serverless being the greatest example. Cloud VMs are neat and all, but in a sense it's the worst of both worlds.

I'm not saying there isn't value in the cloud. I'm saying most people are overpaying. BigCloud didn't solve this any more than paying employees in general to solve your problems. With the era of free money coming to the end, many companies will come to this realization themselves in any case. BigTech clouds have 50%+ margins - something will give eventually.

Any serious company is still going to have to pay oncallers, and admins with or without cloud. It's not like using the cloud absolves you from maintenance (pretty much every company with a valuation more than 100 million has an infrastructure team, and uses the cloud. So clearly the cloud doesn't mean you don't have to maintain your own infrastructure). And now we come back to my point - why isn't the fleet easier to maintain to begin with? I'm not even talking about bare-metal necessarily. Say you use VMs. Still a huge hassle.


> BigTech clouds have 50%+ margins - something will give eventually.

Yes, the oligopoly, and the vendor lock-in.

Certainly won't be expecting the "advantage" of being able to outsource your sysadmins to somewhere else at the drop of a hat to be the thing that gives.


How often do people actually migrate clouds?

Vendor lock in at this level is not really an issue for most, trying to be agnostic will create more problems and time sinks


> Paying premiums for BigCloud means not paying salaries

You need all the same people to manage the same software stack, regardless of whether the VMs sit on hardware you own or at an AWS rack. Nearly all the work is on the software side.

The physical bit of maintaining the boxes is a minimal percentage of the work. At various startups we usually didn't have anyone hired to do this work because there wasn't enough to justify a person. Most simple changes (like pull and replace a hot-plug drive) can be done by the colo personnel and for larger work someone would drive to the colo maybe once a month. You only start to need dedicated hardware maintenance personnel at a very large scale.

The premium you're paying for cpu and bandwidth at your BigCloud is so enormous that it'll easily pay many more salaries than the people you need even if you reach the scale of needing dedicated hardware people.


Yep, in the old days (almost 20 years ago now), I worked at a startup with its own racks. We'd go up there once a month to swap some drives. In my 5 years there, I had to use remote hands exactly twice to reboot a machine. A few other times, one of the servers froze but we were able to reboot it with ILO. This was 3 racks, about 30 machines spread between them (definitely a bunch of empty space, too.)


Happy you, I was in the same situation but I was in the server room daily. Practically every day something went wrong.


I am genuinely curious about both the software and hardware stack if you were there every day due to something going wrong.

1. First of all, a good bare metal setup leaves much of the software fleet management able to be done remotely. So if you were there because of that then clearly this wasn't done right. I'm talking about SSH access at a bare minimum, and ideally out of band access as well.

2. I'd say the things that go wrong in a data center are plentiful, but should be predictable. Power, network or cooling related issues means you picked the wrong site or your vendor screwed you. That leaves us to the actual hardware. Sourcing good hardware is obviously critical. Modern Dell and HP enterprise machines should be giving you very, very reliable hardware assuming the beforementioned concerns are addressed. It is true that disks fail, and sometimes memory, but if you were there literally every day there is just a critical failure in your setup somewhere.

Even if you had a data center with literally a thousand machines. An uncorrelated failure happening every day is so unlikely that it's not really worth mentioning. Correlate failures could certainly happen. Bad batch of disks, etc, but still shouldn't result in you being there literally every day. I could imagine a series of bad days though, sure.


Either you had thousands of machines (at which point is totally makes sense to hire a dedicated person with all the money you're saving) or the site had environmental problems (dirty power, bad cooling, etc?) causing the issues.

Hardware is unbelievably reliable, systems normally run for many years without any attention.


How many systems did you have? What sorts of problems?


You are right, but this was also solved by "dedicated server" folks years before cloud was a thing.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: