Hacker Newsnew | past | comments | ask | show | jobs | submit | muskmusk's commentslogin

Jesus.

This is starting to feel pretty **ing exponential.


I don't think anyone can consider it a success just yet. But someone just injected 2 billion dollars into the enterprise. They are probably given a lot more information than you and I are. Maybe those people are dead wrong and they lose all their money, but maybe not.

With the public information there is neither a point in positivity not negativity. It would be speculation either way. A question that could be worth pondering is what exactly the investors are shown? What would you need to see to put all your money in thinking machines?


The answer was right in Mira's resignation from OpenAI [1]

“I want to create the time and space to do my own exploration.”

To create time and space and even able to explore it is far beyond cutting edge physics of today. If she is able to do it in just few billion dollars it is investment worth making.

1. https://www.wired.com/story/mira-murati-thinking-machines-la...


I love it!

Finally some clear thinking on a very important topic.


This is a surprisingly good idea. The model vs model is fun, but not really that useful.

But this could be a legitimate way to design apps in general if you could tell the models what you liked and didn't like.


yes! that is the hope — /play is our first attempt at building out utility, would love your feedback and will ship hard to make it happen!


I guess masks are completely off now. We can see who sells out to the highest bidder and who won't sell because they care more about the mission.


"everything evolves until it becomes an operating system"


Or at least until it contains an "ad hoc, informally-specified, bug-ridden, slow implementation of half of Common Lisp"


I an not sure these fairly one dimensional views are useful. Completely agree with the premises that it is the managers job to know what their people are doing. I have also worked at companies where the managers weren't doing their job. I can also completely see how an overfocus on metrics kills an otherwise good company.

I don't agree with the conclusion though: metrics are bad.

In software, in theory, you don't need metrics. You just write good code that does what it needs to do. Add tests to prove it, you are good. In reality though everyone agrees that metrics are a good thing to have. Why do we need them? Complexity and sometimes we just make mistakes. We could write software that is very close to bug free, but then we move at aerospace speed. Not always optimal.

Software metrics come with all the pitfalls of management metrics like over focus on metrics and not seeing the forest for the trees. Yet we still find the metrics useful.

Managers face many of the same problems, so they ask for metrics. I don't think that is bad. Being bad at using metrics is bad.


Blog post is about someone approaching author on "individual employee productivity tooling".

So it is not "all metrics are bad", it is "individual employee productivity tracking is bad".

On team level and across teams having trends and looking at trends before and after some change is helpful and can be done only when you have metrics of course.


It's a free world. Nothing stops you from applying their findings to bigger datasets. It would be a valuable contribution.


Yes.


Friend, the creator of this new progress is a machine learning PhD with a decade of experience in pushing machine learning forward. He knows a lot of math too. Maybe there is a chance that he too can tell the difference between a meaningless advance and an important one?


That is as pure an example of the fallacy of argument from authority[1] as I have ever seen especially when you consider that any nuance in the supposed letter from the researchers to the board will have been lost in the translation from "sources" to the journalist to the article.

[1] https://en.wikipedia.org/wiki/Argument_from_authority


That fallacy's existence alone doesn't discount anything (nor have you shown it's applicable here), otherwise we'd throw out the entire idea of authorities and we'd be in trouble


Authorities are useful within a context. Appealing to authority is not an argument. At most, it is an heuristic.

_Using_ this fallacy in an argument invalidates the argument (or shows it did not exist in the first place)


When the person arguing uses their own authority (job, education) to give their answer relevance, then stating that the authority of another person is greater (job, education) to give that person's answer preeminence is valid.


I am neither a mathematician or LLM creator but I do know how to evaluate interesting tech claims.

The absolute best case scenario for a new technology is that it when it seems like a toy for nerds, and doesn't outperform anything we have today, but the scaling path is clear.

Its problems just won't matter if it does that one thing with scaling. The web is a pretty good hypermedia platform, but a disastrously bad platform for most other computer applications. Nevertheless the scaling of URIs and internet protocols have caused us to reorganize our lives around it. And then if there really are unsolvable problems with the platform they just get offloaded onto users. Passwords? Privacy? Your problem now. Surely you know to use a password manager?

I think this new wave of AI is going to be like that. If they never solve the hallucination/confabulation issue, it's just going to become your problem. If they never really gain insight, it's going to become your problem to instruct them carefully. Your peers will chide for not using a robust AI-guardrail thing or not learning the basics of prompt engineering like all the kids do instinctively these days.


How on earth could you evaluate the scaling path with too little information. That's my point. You can't possibly know that a technology can solve a given kind of problem if it can only so far solve a completely different kind of problem which is largely unrelated!

Saying that performance on grade-school problems is predictive of performance on complex reasoning tasks (including theorem proving) is like saying that a new kind of mechanical engine that has 90% efficiency can be scaled 10x.

These kind of scaling claims drive investment, I get it. But to someone who understands (and is actually working on) the actual problem that needs solving, this kind of claim is perfectly transparent!


Any claims of objective, quantitative measurements of "scaling" in LLMs is voodoo snake oil when measured against some benchmarks consisting of "which questions does it answer correctly". Any machine learning PhD will admit this, albeit only in a quiet corner of a noisy bar after a few more drinks than is advisable when they're earning money from companies who claim scaling wins on such benchmarks.


For the current generative AI wave, this is how I understand it:

1. The scaling path is decreased val/test loss during training.

2. We have seen multiples times that large decreases in this loss have resulted in very impressive improvements in model capability across a diverse set of tasks (e.g. gpt-1 through gpt-4, and many other examples).

3. By now, there is tons of robust data demonstrating really nice relationships between model size, quantity of data, length of training, quality of data, etc and decreased loss. Evidence keeps building that most multi-billion param LLMs are probably undertrained, perhaps significantly so.

4. Ergo, we should expect continued capability improvement with continued scaling. Make a bigger model, get more data, get higher data quality, and/or train for longer and we will see improved capabilities. The graphs demand that it is so.

---

This is the fundamental scaling hypothesis that labs like OpenAI and Anthropic have been operating off of for the past 5+ years. They looked at the early versions of the curves mentioned above, extended the lines, and said, "Huh... These lines are so sharp. Why wouldn't it keep going? It seems like it would."

And they were right. The scaling curves may break at some point. But they don't show indications of that yet.

Lastly, all of this is largely just taking existing model architectures and scaling up. Neural nets are a very young technology. There will be better architectures in the future.


We're at the point now where the harder problem is obtaining the high quality data you need for the initial training in sufficient quantities.


These European efforts to create competitive LLMs need to know that.


I don't think they will go anywhere. Europe doesn't have the ruthlessness required to compete in such an arena, it would need far more unification first before that could happen. And we're only drifting further apart it seems.


I didn’t say “certain success”, I said “interesting”


Honestly, OpenAI seem more like a cult that a company to me.

The hyperbole that surrounds them fits the mould nicely.


They did build the most advanced LLM tool, though.


Maybe it takes a cult


But he also has the incentive to exaggerate the AI's ability.

The whole idea of double-blind test (and really, the whole scientific methodology) is based on one simple thing: even the most experienced and informed professionals can be comfortably wrong.

We'll only know when we see it. Or at least when several independent research groups see it.


> even the most experienced and informed professionals can be comfortably wrong

That's the human hallucination problem. In science it's a very difficult issue to deal with, only in hindsight you can tell which papers from a given period were the good ones. It takes a whole scientific community to come up with the truth, and sometimes we fail.


No. It takes just one person to come up with the truth. It then can takes ages to convince the "scientific community".


Well, one person will usually add a tiny bit of detail to the "truth". It's still a collective task.


I don't think so. The truth is advanced by individuals, not by the collective. The collective is usually wrong about things as long as they possibly can be. Usually the collective first has too die before it accepts the truth.


I thought (and could be wrong) that all of these concerns are based on a very low probability of a very bad outcome.

So: we might be close to a breakthrough, that breakthrough could get out of hand, then it could kill a billion+ people.


> I thought (and could be wrong) that all of these concerns are based on a very low probability of a very bad outcome.

Among knowledgeable people who have concerns in the first place, I'd say giving the probability of a very bad outcome of cumulative advances as "very low" is a fringe position. It seems to vary more between "significant" and "close to unity".

There are some knowledgeable people like Yann LeCun who have no concerns whatsoever but they seem singularly bad at communicating why this would be a rational position to take.


Given how dismissive LeCun is of the capabilities of SotA models, I think he thinks the state of the art is very far from human, and will never be human-like.

Myself, I think I count as a massive optimist, as my P(doom) is only about 15% — basically the same as Russian Roulette — half of which is humans using AI to do bad things directly.


Unlikely. We'll know when OpenAI has declared itself ruler of the new world, imposes martial law, and takes over.


Why would you ever know? Why would the singularity reveal itself in such an obvious way(until it's too late to stop it)?


Also, wbhart is referring to publicly released LLMs, while the OpenAI researchers are most likely referring to an un-released in-research LLM.


Sure... but that machine learning PhD has a vested interest in being optimistically biased in his observations.


Ah finally the engineers approach to the news. I'm not sure why we have to have hot takes, instead of dissecting the news and trying to tease out the how.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: