> You can make a safe AI today, but what happens when the next person is managing things?
The point of safe superintelligence, and presumably the goal of SSI Inc., is that there won't be a next (biological) person managing things afterwards. At least none who could do anything to build a competing unsafe SAI. We're not talking about the banal definition of "safety" here. If the first superintelligence has any reasonable goal system, its first plan of action is almost inevitably going to be to start self-improving fast enough to attain a decisive head start against any potential competitors.
Trouble is, in practice what you would need to do might be “turn off all of Google’s datacenters”. Or perhaps the thing manages to secure compute in multiple clouds (which is what I’d do if I woke up as an entity running on a single DC with a big red power button on it).
The blast radius of such decisions are large enough that this option is not trivial as you suggest.
a) after you create the superintelligence is likely too late. You seem to think that inventing superintelligence means that we somehow understand what we created, but note that we have no idea how a simple LLM works, let alone an ASI that is presumably 5-10 OOM more complex. You are unlikely to be able to control a thing that is way smarter than you, the safest option is to steer the nature of that thing before it comes into being (or, don’t build it at all). Note that we currently don’t know how to do this, it’s what Ilya is working on. The approach from OpenAI is roughly to create ASI and then hope it’s friendly.
b) except that is not how these things go in the real world. What actually happens is that initially it’s just a risk of the agent going rogue, the CEO weighs the multi-billion dollar cost vs. some small-seeming probability of disaster and decides to keep the company running until the threat is extremely clear, which in many scenarios is too late.
(For a recent example, consider the point in the spread of Covid where a lockdown could have prevented the disease from spreading; likely somewhere around tens to hundreds of cases, well before the true risk was quantified, and therefore drastic action was not justified to those that could have pressed the metaphorical red button).
> Having arms and legs is going to be a significant benefit for some time yet
I am also of this opinion.
However I also think that the magic shutdown button needs to be protected against terrorists and ne'er-do-wells, so is consequently guarded by arms and legs that belong to a power structure.
If the shutdown-worthy activity of the evil AI can serve the interests of the power structure preferentially, those arms and legs will also be motivated to prevent the rest of us from intervening.
So I don't worry about AI at all. I do worry about humans, and if AI is an amplifier or enabler of human nature, then there is valid worry, I think.
Where can I find the red button that shuts down all Microsoft data centers, all Amazon datacenters, all Yandex datacenters and all Baidu datacenters at the same time? Oh, there isn't one? Sorry, your superintelligence is in another castle.
It's been more than a decade now since we first saw botnets based on stealing AWS credentials and running arbitrary code on them (e.g. for crypto mining) - once an actual AI starts duplicating itself in this manner, where's the big red button that turns off every single cloud instance in the world?
Is that really "a lot of assumptions" that a piece of software can clone itself? We've been cloning and porting software from system to system for over 70 years (ENIAC was released in 1946 and some of its programs were adapted for use in EDVAC in 1951) - why would it be a problem for a "super intelligence"?
And even if it was originally designed to run on some really unique ASIC hardware, by the Church–Turing thesis it can be emulated on any other hardware. And again, if it's a "super intelligence", it should be at least as good at porting itself as human engineers have been for the three generations.
A "state of the art" system would almost by definition be running on special and expensive hardware. But I have llama3 running on my laptop, and it would have been considered state of the art less than 2 years ago.
A related point to consider is that a superintelligence should be considered a better coder than us, so the risk isn't only directly from it "copying" itself, but also from it "spawning" and spreading other, more optimized (in terms of resources utilization) software that would advance its goals.
This is why I think it’s more important we give AI agents the ability to use human surrogates. Arms and legs win but can be controlled with the right incentives
> there won't be a next (biological) person managing things afterwards. At least none who could do anything to build a competing unsafe SAI
This pitch has Biblical/Evangelical resonance, in case anyone wants to try that fundraising route [1]. ("I'm just running things until the Good Guy takes over" is almost a monarchic trope.)
The point of safe superintelligence, and presumably the goal of SSI Inc., is that there won't be a next (biological) person managing things afterwards. At least none who could do anything to build a competing unsafe SAI. We're not talking about the banal definition of "safety" here. If the first superintelligence has any reasonable goal system, its first plan of action is almost inevitably going to be to start self-improving fast enough to attain a decisive head start against any potential competitors.