UnionFaps
AIExplained from patreon
AIExplained patreon

Reflections on Sam Altman’s Recent Expectation-Setting on GPT-5 

🕑 Added 2024-04-28 13:38:44 +0000 UTC

Comments

Clay Farris Naff

I appreciate your taking the time to reply and share further thoughts. I did not mean to imply that scientific "truths" are merely the consensus opinion of experts. To be science, they must at least correspond to experimental results or observed reality. However, as demonstrated by gravity and QM, the explanation of a scientific truth can vary widely, and it evolves over time as new data becomes available. The best way of characterizing this, IMHO, is Asimov's "relativity of wrong." We never get to The Truth, but we do get progressively closer. Hawking, in his own way, concurred. Cheers, Clay

Kofi Baah

> Self-verification/critique doesn’t work for current model, though they haven’t trained for it specifically; unclear if it is possible to train for, but believes it will be; we will figure it out in within the next few model versions Why not train a verification model separately to the generation model? Are there approaches which explore this?

Kofi Baah

I'd disagree with the idea that scientific truths are rooted in the consensus of experts. Scientific truths are good explanations of physical reality that successfully solve a problem. My intuition is that yes, verification will lead to AI being a reliable source of truth. Not necessarily self-verification in the tree-of-thought sense, but approaches that test a model's conjectures against physical reality. See: Philip's "AI Conquers Gravity" video https://www.youtube.com/watch?v=d5mdW1yPXIg Funsearch https://deepmind.google/discover/blog/funsearch-making-new-discoveries-in-mathematical-sciences-using-large-language-models/

Feitian Li

Note Altman's wording "economic models will break." This is much more extreme than merely mass unemployment. For most of human history society consisted entirely of a few wealthy and powerful elites and a mass of underemployed and violent underclass, yet the economic and political model didn't break despite that. To suggest a complete break of economic model which is usually the last link of the society to breakdown would imply an unimaginable amount of displacement ina very short amount of time, and most importantly it must disempower the elite as much as it does everybody else.

Feitian Li

OpenAI employees seem to be extremely confident agi will arrive in a reasonably short amount of time. This is increasingly shocking to me, because as I reckon this is still a relatively open field with a mostly publicly available literature, and nothing has suggested definitively yet that the rest of the unknowns can be figured out.

GGuy

I guess the more specific questions are: Is the training problem solved with synthetic data? Does it even matter if the synthetic data is wrong? Is the RAG problem solved with 10m+ context windows? RAG = Insert all the things until 9m tokens is reached. Is that easier and more effective then the RAG we are going today with smaller context limits?

Jan Wilczynski

Isn't it obvious? Like we already know that domain specific training and in-context learning drastically improves the outcomes and big part of the battle now is about the scaffolding to effectively achieve it. In training the problem is data, in context the problem is that RAGs don't work yet the way we want them to ;D

Kemi

"Dario Amodei, CEO of Anthropic - ‘Nothing truly insane happens in 2024’, and that most of the impressive stuff is ‘2026 onwards’." I'm sure this is a well considered comment from Mr Amodei, but I do also wonder if its subject to perspective? The area under the "truly insane" curve feels like it could still encompass significant change. In particular my mind wanders to the performance of agentic GPT 3.5 systems vs GPT-4 (as recently discussed by Andrew Ng in a series from Sequoia Capital IIRC). Can't help but wonder what such clever architectural systems built around GPT-5 might end up being able to do. Mind you the labs probably have such systems behind closed doors already, so perhaps we'll all start to learn what those capabilities are as research discovers them, while the labs look on from their position some number of months in the future knowing what comes next :p

GGuy

"In context" of recent papers and the phi-based architectures I wonder if developers are considering a new approach: providing models with all necessary in-context information during both training and inference. Instead of relying on encoded knowledge from training data, this method would supply relevant information as part of the input context for each response generation. I wonder if this approach could potentially improve the model's reasoning capabilities and lead to more accurate, context-aware responses.

Sean Gallagher

Understandable that everyone has focused on Sam's comments lately. But I think Brad Lightcap's contributions in the VC20 interview should get more air time because he is focused on where the rubber hits the road with enterprises. He says, “enterprises have a very natural desire to want to throw the technology into a business process with the pure intent of driving a very quantifiable ROI. ‘I want to take AI and throw it at a very specific process in supply chain management and cut 20% of my spend’.” He thinks leaders, “criminally underrate” how much return you really get by just providing people access to the technology. You can’t quite quantify an exact ROI from GenAI. This is precisely my experience in speaking to leaders. They want to see LLMs as a technology that can solve a specific business problem and hone in on that. Rather a mindset shift is required to see GenAI as a virtual coworker - or army of virtual coworkers. For Moderna to launch 15 new products over 5 years without AI would require 100,000 people, but with technology like ChatGPT they will be able to do it with less than 6,000 workers. What's the ROI on that?

adfaklsdjf

Speaking as someone in the USA, I felt instantly gloomy when I read the comparison to handling the pandemic... I seriously hope we handle AGI/ASI far better than we did the pandemic, but it might be reasonable to use that benchmark to set expectations.. :(

Joe Marler

Philip, as always you've provided incisive and well-reasoned analysis. I use GPT-4 Turbo and Claude 3.0 Opus every day. I'm very glad to have them but they have limits I'm not sure can be solved using self-supervised learning -- no matter how smart the model is. E.g, I frequently ask them syntax questions about Python, ffmpeg, zsh, etc. They often make mistakes since there are many versions of those with subtly different syntax. Even if I tell them up front what versions I'm using, the errors still happen. It finally occurred to me -- that version-specific metadata was not captured in the unsupervised training. It further occurred to me that other areas are likely affected by this—legal advice, medical advice, technical troubleshooting, etc. The core LLM principle is unsupervised learning, which hopes that bulk data, scalability, and mysterious emergent behavior will enable further progress. Yet in my limited experience, it appears there may be areas where progress may be constrained due to the lack of metadata in the corpus. I think progress will still happen, but the LLMs may exhibit more "spikey" capability across knowledge domains. It will be interesting to see whether GPT-5 and the other frontier models improve in these areas.

Arek Stryjski

I'm wondering how this statement applies to It and sturtups. Will the model brake? If you build new company on the GTP-4 or 5, will it be whipped out by 6 or 7?

Edward Huff

So either the data-set optimization gains are overhyped or he’s lying. He does have every incentive to lie but so far he’s been reasonably honest about internal capabilities. But if you are near to AGI, the move is to lie until you have a decisive strategic advantage and become God Emperor.

Cyril Sadovsky

From the tone of your commentary and Sam Altman's remarks, it's clear that the market is beginning to move past the initial GenAI hype. I've felt this myself. As a developer, I recently stopped using GitHub Copilot because I found it was actually slowing me down. I've noticed other developers experiencing the same trend. However, this shift could be beneficial. It provides an opportunity for innovators and engineers to apply this new technology to real-world business scenarios, though it will require time. Thank you for your commentary, Philipp.

Taunger

Adaptive compute is going to be the most important thing going forward, I expect. It’s so strange that people expect these models to do everything right at the first thought. We don’t expect that from humans. The GPT models can already think. They just need time to think.

Steve DeMoss

Also, Amara's law states "Technology tends to be overestimated in the short term and underestimated in the long term." Of course it depends on your definition of what is short and long but this may still be applicable here.

Steve DeMoss

"Current economic models will break." Really would love more context here as that's potentially an incredibly dramatic statement. Models for the AI industry? Macro economic models? The global economy? Maybe just understanding the question that led to that comment would help.

Clay Farris Naff

I've been mulling AI epistemology lately. After an Intro to AI presentation I gave, an audience member asked me if AI will be able to cut through all the misinformation in the infosphere and tell us what's true. I gave a very cautious (and inadequate) answer. It made me think about how even mathematical truths only hold for a given set of axioms, and how scientific truths are always tentative and rooted in the consensus of experts, which evolves over time. As for political or historical truths? Who can say? But I am curious to know if you think that verification procedures will lead, eventually, to AI as a reliable source of truth.

Sean Betts

I’m still wondering whether we’ll see a GPT-4.5 before GPT-5, based on comments Sam made in his podcast with Lex Fridman a month or so ago… he reflected on whether OpenAI should have more iterative release cycles as opposed to big annual releases to help the public adapt more. Whether it’s GPT-4.5 or GPT-5 I think the things that will have the biggest impact in the short term will be the engineering OpenAI do around the models, not necessarily the models themselves. The things you mention like video avatars, real-time conversations, personalisation and the start of agent-like behaviour will have much bigger real-world impacts than the overall intelligence gains we’re likely to see between generations of models at the moment. However, I do think we will see a breakthrough in model intelligence at some point (maybe 2026/27) but it will be based on a new type of architecture and training approach, probably along the lines of the objective-driven AI approach Yann LeCun has been proposing.

Eddine Maiza

Thanks for sharing! Could you share how you are using AI to help you with your work? (Eg to parse through openAI employees comments on X?)

Machiel Reyneke

"- The current economic models will break -Capitalism will have to change" I realise that most, if not everyone here would intuitively agree with that statement. I do. I'd be surprised if many people at research labs disagree. I guess the question is: "how bad will it have to get before the system is forced to change?" In a way wouldn't it be better if an ASI drops unexpectedly and we need to adapt ASAP (similar to in a pandemic)? As opposed to the incremental change where the frustration of people build as more and more jobs are displaced and more and more people are disgruntled as new models become more powerful.

Robert Gomez-Reino

Wow, I totally missed the one about GPT4-6-8... I feel that GPT4 is already helping me in multiple tasks which I don't consider basic at all. I wonder if this means like ChatGPT or really what one can build around GPTX. If really GPT6 will be JUST saying "this is helping me as a general-purpose tool..." then we (at least me) are up to disappointment regarding GPT5. The quality I am most interested on is reasoning power, the rest can be built around with today's tools (ok, some decent inference speed to enable real time use cases might be required too) and Sama promised a leap on that one. I find quite frustrating the ambiguity around that topic: "you need to plan and build big or we will steamroll you, but I cannot really tell you how different the reasoning of this next model is". Could it be that it is really difficult to express how smart the models are from now on? So we can only say they will be "way smarter" or "significant smarter" (this was used by Sama for GPT4 April release btw...). If GPT5 doesn't deliver in reasoning as promised (same as from 3.5 to 4) then I will feel I was lied to.

Trenton Dambrowitz

Sad only one question went through to Sam, but I think it was the most important one. When you think about it from a business’s perspective, in-context learning is probably the most important thing to get right. If I’m trying to teach an agent how to order parts for a car we’re fixing I don’t care how well it does on HellaSwag or MMLU, it needs to understand our processes and follow them properly! I would absolutely love to see how many shot prompting, GPT-5 level reasoning and In-context learning, and SmartGPT combine on multi-modal tasks that the model wasn’t originally designed for.

Anouar Mansour

So slow incremental progress is going to be the way for the next couple years

David Shapiro

Great analysis. I listened to the same comments for clues from Sam. I'm recalling how, just a few months ago, we were all debating about whether or not GPT-4 would be disappointing, and in reflection, yes, I think Sam was right in forecasting that GPT-4 has been a bit disappointing. What I'm MOST curious about is what theory of mind or intelligence they are using to forecast what abilities and disruptive ability these models have, or if it's just intuition from getting early access.

Steven

Thanks for the update. I’m convinced that the short term investment is/was very high and could crash a bit because the money is just rolling in.


More Creators