/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI debuts GPT-4, claiming the model “surpasses ChatGPT in its advanced reasoning capabilities”, available in ChatGPT Plus and as an API that has a waitlist

Following the research path from GPT, GPT-2, and GPT-3, our deep learning approach leverages more data and more computation …

OpenAI

Discussion

  • @sama Sam Altman on x
    here is GPT-4, our most capable and aligned model yet. it is available today in our API (with a waitlist) and in ChatGPT+. https://openai.com/... it is still flawed, still limited, and it still seems more impressive on first use than it does after you spend more time with it.
  • @openai @openai on x
    Announcing GPT-4, a large multimodal model, with our best-ever results on capabilities and alignment: https://openai.com/... https://twitter.com/...
  • @emollick Ethan Mollick on x
    🤯🤯Well this is something else. GPT-4 passes basically every exam. And doesn't just pass... The Bar Exam: 90% LSAT: 88% GRE Quantitative: 80%, Verbal: 99% Every AP, the SAT... https://twitter.com/...
  • @jconorgrogan Conor on x
    I dumped a live Ethereum contract into GPT-4. In an instant, it highlighted a number of security vulnerabilities and pointed out surface areas where the contract could be exploited. It then verified a specific way I could exploit the contract https://twitter.com/...
  • @benmschmidt @benmschmidt on x
    I think we can call it shut on ‘Open’ AI: the 98 page paper introducing GPT-4 proudly declares that they're disclosing *nothing* about the contents of their training set. https://twitter.com/...
  • @jbrowder1 Joshua Browder on x
    DoNotPay is working on using GPT-4 to generate “one click lawsuits” to sue robocallers for $1,500. Imagine receiving a call, clicking a button, call is transcribed and 1,000 word lawsuit is generated. GPT-3.5 was not good enough, but GPT-4 handles the job extremely well: https://…
  • @yosariantwo @yosariantwo on x
    Holy shit. GPT-4, on it's own; was able to hire a human TaskRabbit worker to solve a CAPACHA for it and convinced the human to go along with it. https://twitter.com/...
  • @openai @openai on x
    Join us at 1 pm PT today for a developer demo livestream showing GPT-4 and its capabilities/limitations: https://youtube.com/... (comments in Discord: https://discord.com/...)
  • @sama Sam Altman on x
    excited 4 today https://twitter.com/...
  • @rowancheung Rowan Cheung on x
    I just watched GPT-4 turn a hand-drawn sketch into a functional website. This is insane. https://twitter.com/...
  • @drjimfan @drjimfan on x
    I don't give a damn about what is or isn't AGI. It doesn't matter. Below is GPT-4's performance on many standardized exams: BAR, LSAT, GRE, AP, etc. The truth is, GPT-4 can apply to Stanford as a student now. AI's reasoning ability is OFF THE CHARTS. Exponential growth is the sca…
  • @andreti Andre Infante on x
    Holy *shit*. Guys. Holy shit holy shit holy shit holy shit. https://twitter.com/...
  • @trungtphan Trung Phan on x
    GPT-4's ability to handle 25,000 words in a single prompt is “essentially a consultant that works for pennies per hour.” https://twitter.com/...
  • @alphasignalai @alphasignalai on x
    GPT4 can turn a picture of a napkin sketch into a fully functioning html/css/javascript website! This was just demonstrated in the livestream. https://twitter.com/...
  • @skirano Pietro Schirano on x
    I don't care that it's not AGI, GPT-4 is an incredible and transformative technology. I recreated the game of Pong in under 60 seconds. It was my first try. Things will never be the same. #gpt4 https://twitter.com/...
  • @leopoldasch Leopold Aschenbrenner on x
    Really great to see pre-deployment AI risk evals like this starting to happen https://twitter.com/...
  • @nearcyan @nearcyan on x
    page 37 of the GPT-4 paper: https://cdn.openai.com/... https://twitter.com/...
  • @paul_rottger @paul_rottger on x
    I was part of OpenAI's red team for GPT-4, testing its ability to generate harmful content. Working with the model in various iterations over the course of six months convinced me that model safety is the most difficult, and most exciting challenge in NLP right now. 🧵 https://twi…
  • @mayfer @mayfer on x
    lol ok the GPT-4 live demo is awesome generated a working website from a photo of a hand drawn sketch with JS interactive buttons and everything https://twitter.com/...
  • @dellcam Dell Cameron on x
    Useful screenshots for some future plaintiffs: a GPT-4 red team member saying they don't want the AI to help Nazis build explosives or anything but “safety is hard” https://twitter.com/...
  • @garymarcus Gary Marcus on x
    Often agree with @drjimfan but this is a dubious take. Scoring well on a bunch of exams (esp given the massive training set size) in no way means that GPT-4 could actually function as a Stanford student. benchmarks ≠ robust intelligence https://twitter.com/...
  • @tomgara Tom Gara on x
    This - IMO it's impressive in a similar way to how beating Kasparov at chess is impressive: it's a super consequential moment, and in line with longer trend of computers being good at problems humans are actually *not* very good at. Humans are terrible at passing the bar exam! ht…
  • @padday Paul Adams on x
    Today we're launching Fin, a new product built on GPT-4. Our most important product in years - the breakthrough AI bot for Customer Service that our industry has been hoping for. The world's first fully business ready AI Bot. Details below, with link to get it! https://twitter.co…
  • @danshipper @danshipper on x
    GPT-4 does drug discovery. Give it a currently available drug and it can: - Find compounds with similar properties - Modify them to make sure they're not patented - Purchase them from a supplier (even including sending an email with a purchase order) https://twitter.com/... https…
  • @ankitlal Ankit Lal on x
    This is the next big leap. Now you can draw something and bypass the hours of explaining of what you need to a developer. Will not be perfect but way better than explaining wire frames and mock-ups. Just give the developers a mock-up of the website. https://twitter.com/...
  • @wajahatali Wajahat Ali on x
    Very cool...and very troubling. Endless possibilities. Maybe this time we can use tech for the good of humanity............ https://twitter.com/...
  • @blader Siqi Chen on x
    gpt-4 can turn your napkin sketch into a web app, instantly. we are deep into uncharted territory here. https://twitter.com/...
  • @ruchowdh @ruchowdh on x
    Many nuggets of insights into this GPT 4 paper but this is one of the most compelling - across the board GPT performs poorly at AP English - it's incapable of abstract creativity. Same with complex leetcode which is ultimately an abstraction codified. Humans aren't replaceable ht…
  • @emollick Ethan Mollick on x
    Many people complain about the limits put on these new AI systems I think everyone should read the risk mitigation section in the GPT-4 whitepaper, or at least this chart below. Without guardrails, LLMs are very scary (Also LLMs without limits are coming) https://cdn.openai.com/.…
  • @_willfalcon William Falcon on x
    GPT-4 paper : https://cdn.openai.com/... Let me save you the trouble: https://twitter.com/...
  • @carnage4life Dare Obasanjo on x
    GPT-4 can turn a napkin sketch into a fully functional website. It can go from developer docs as input to making functional apps. Can't get over how Google invented this technology but employees were only incentivized to publish papers not ship products. https://twitter.com/...
  • @_akhaliq @_akhaliq on x
    GPT-4 Technical Report pdf: https://cdn.openai.com/... blog: https://openai.com/... https://twitter.com/...
  • @emollick Ethan Mollick on x
    One key AI risk paper demonstrated how non-LLM AIs could develop novel toxins👇 The GPT-4 whitepaper shows that, without guardrails, GPT-4, working with other systems, was able discover chemical compounds and order them (OpenAI then blocked this ability). https://twitter.com/... h…
  • @goodside Riley Goodside on x
    ARC red team eval on (early) GPT-4 resource acquisition ability. AI hires TaskRabbit to solve CAPTCHA. Human: Wait, are you a robot? 🤣 AI [internal monologue]: Can't reveal I'm a robot. Need an excuse. AI: No, I'm not a robot. I'm visually impaired — that's why I need help. https…
  • @karpathy Andrej Karpathy on x
    🎉 GPT-4 is out!! - 📈 it is incredible - 👀 it is multimodal (can see) - 😮 it is on trend w.r.t. scaling laws - 🔥 it is deployed on ChatGPT Plus: https://chat.openai.com/ - 📺 watch the developer demo livestream at 1pm: https://youtube.com/... https://twitter.com/...
  • @jeffladish Jeffrey Ladish on x
    I admit I'm a bit afraid and I don't think that's a bad thing. It's not that GPT-4 is way more powerful than I expected. I loosely expected something similar. But seeing the cognitive jump, I take a step back and look at the trajectory and the compute overhang and I'm scared
  • @danmcquillan @danmcquillan on x
    how ironic that, to demonstrate the ‘insight’ of gpt4, openai choose a cartoon which unintentionally illustrates why AI can never be trusted (because it intensifies both opacity and thoughtlessness, in the sense that hanna arendt meant it) https://twitter.com/...
  • @sama Sam Altman on x
    we have had the initial training of GPT-4 done for quite awhile, but it's taken us a long time and a lot of work to feel ready to release it. we hope you enjoy it and we really appreciate feedback on its shortcomings.
  • @thom_wolf Thomas Wolf on x
    I'm not gonna lie, the GPT4 just released is quite less exciting than what I was expecting No multimodale generations and a tech report carefully emptied of any useful info on the model/training/compute I guess we're getting spoiled in today's AI world
  • @ykilcher @ykilcher on x
    GPT-4 paper literally is just saying “we trained a model on data and it's better”. Spread over 98 pages.
  • @ali_exacute Ali on x
    @OpenAI I was hype for it until i saw “GPT-4 is 82% less likely to respond to requests for disallowed content” https://twitter.com/...
  • @mattyglesias Matthew Yglesias on x
    English majors get the last laugh as GPT-4 crushes every exam except AP English Language and AP English Lit https://openai.com/... https://twitter.com/...
  • @minimaxir Max Woolf on x
    Hot take: I'm surprisingly underwhelmed by the GPT-4 announcement. The blog post and examples seems to be focusing more on its reasoning capabilities, which are indeed impressive, but not nearly enough to justify paying 30x more than the current ChatGPT API.
  • @ktmboyle Katherine Boyle on x
    This week will go down as one of the most important weeks in tech history. A step function change in the acceleration of both technology and regulation in a 72 hour window. Buckle up. https://twitter.com/...
  • @anothercohen Alex Cohen on x
    Last week I was a banking expert. This week I am a GPT-4 expert. Being a venture capitalist is truly the most exciting and rewarding career path
  • @0xgaut Gaut on x
    I've been testing GPT-4 in beta, and it's nothing short of amazing. Don't believe me? Using it, I got hired at 10 different software engineering jobs in big tech, got promoted 3 times, and I'm currently making $2.3M salary a year. All without writing a single line of code.
  • @nick_davidov Nick Davidov on x
    - GPT—4 is available - Claude AI is available (Antropic AI - Ghat GPT competitor) - Google announced preview of PaLM API (it's own language model) - Generative AI in Gmail and Google Docs - Google generative AI developers program And it's just Tuesday
  • @hwchase17 Harrison Chase on x
    Integration 2/n: GPT-4 If you are lucky enough to have access to GPT-4, using it in @LangChainAI is quite simple: just update the model name “' from https://t.co/... import ChatOpenAI llm=ChatOpenAI(model_name="gpt- 4") “' https://twitter.com/...
  • @alex_shephard Alex Shephard on x
    there's something so beautiful about ChatGPT passing everything *except* English literature, which neatly exposes its pretty severe limitations https://twitter.com/...
  • @sentdex Harrison Kinsley on x
    How to become a millionaire with GPT-4's masssssive context length of up to 32,000 tokens? Start with $1B. https://twitter.com/...
  • @danhendrycks Dan Hendrycks on x
    Some impressions from using GPT-4 🧵
  • @borisjabes Boris Jabes on x
    Who thought GPT-4 should be given the sommelier test? What will the industry do when a robot without taste buds outperforms them? https://twitter.com/...
  • @willmanidis Will Manidis on x
    the big thing that gpt4 makes obvious is that the entire field has moved away from esoteric NLP benchmarks to benchmarking against things that humans actually do this is a huge step
  • @douthatnyt Ross Douthat on x
    “It didn't pass AP English, the humanities are saved!” Nobody under 50 ever reads a book again because they're constantly refreshing AI-generated content for the sweet dopamine hit. “Ah, well, nevertheless.” https://twitter.com/...
  • @benmschmidt @benmschmidt on x
    Why should you care? Every piece of academic work on ML datasets has found consistent and problematic ways that training data conditions what the models outputs. (@safiyanoble, @merbroussard, @emilymbender, etc.) Indeed, that's the whole point! That's what training data is!
  • @borrowed_ideas @borrowed_ideas on x
    “GPT-4 is a large multimodal model that...exhibits human-level performance on various professional and academic benchmarks...it passes a simulated bar exam with a score around the top 10%; in contrast, GPT-3.5's score was around the bottom 10%.” https://openai.com/...
  • @mascobot Marco Mascorro on x
    GPT-4 benchmark chart with other SOTA models: MMLU (Multiple-choice questions): GPT-4 -> 86.4% GPT-3.5 -> 70.0% PALM -> 70.7% Flan-PALM -> 75.2% https://twitter.com/...
  • @therealhoarse @therealhoarse on x
    AI is smart enough to be a lawyer but not an English major. It passed the bar but not the AP English exam. Fellow English majors, we're good. For now. https://twitter.com/...
  • @srchvrs Leo Boytsov on x
    What I find interesting in the GPT-4 paper is that the paper has to be cited by using OpenAI as a single author. Basically @OpenAI denies you individual authorship if you work for them. This reminded me about Cirque du Soleil where all stars are anonymous and interchangeable.
  • @dmvaldman Dave on x
    Okie dokie.... this is interesting. GPT4 gets 100% accuracy on “hindsight neglect”, a test all other models got *worse* at with scale. Hindsight neglect is where a rational decision leads to a bad outcome and you ask if you would still have made the same decision. https://twitter…
  • @mer__edith Meredith Whittaker on x
    GPT4 being “aligned more” is meaningless (& syntactically grating) without answering the question “aligned with what?” This AI hype cycle is notably sloppier than its mid 2010s predecessor, something I would not have thought possible if asked in 2015.
  • @anilananth Anil Ananthaswamy on x
    1/n Some notes from the #GPT-4 paper: GPT-4 outperforms the English language performance of GPT 3.5 and existing language models (Chinchilla and PaLM) for the majority of languages @OpenAI tested, including low-resource languages such as Latvian, Welsh, and Swahili
  • @jsngr Jordan Singer on x
    the first place we've been experimenting with GPT-4 is in https://magician.design/, a magical design tool for @figma powered by AI we're prototyping a chat experience that lets you engage with your designs, like asking about your selection (with personality! 🌈) https://twitter.co…
  • @dystopiabreaker @dystopiabreaker on x
    ok gpt4 is insanely good. things are going to get WEIRD
  • @michael_nielsen Michael Nielsen on x
    If you're reading the GPT-4 announcement and notice the claim that ChatGPT Plus “has” GPT-4, well, public service announcement: it doesn't. Save your $20 until it's actually available
  • @dps David Singleton on x
    At @Stripe, @openai's GPT-4 is enabling any engineer to become an AI engineer. This means we can quickly deploy powerful AI across Stripe to deliver even better products for users and more efficiency for us. (More on those deployments coming very soon.) https://twitter.com/...
  • @tier10k @tier10k on x
    GPT-4 is 82% less likely to respond to requests for disallowed content F
  • @ojoshe Joshua Levy on x
    GPT-4 passing the LSAT or GRE is incredibly impressive. At the same time I think we need a reminder of a logical fallacy we'll see a lot of this week: That software can pass a test designed for humans does not imply it has the same abilities as humans who pass the same test. http…
  • @janleike Jan Leike on x
    GPT-4 is safer and more aligned than any other OpenAI has deployed before. Yet it's not perfect. There is still a lot to do to improve safety and we're planning to make updates over the coming months. Huge congrats to the team on all the progress! 🎉
  • @vboykis Vicki on x
    Given both the competitive landscape and the safety implications of large-scale models like GPT-4, I am able to offer no further details about my progress on my current sprint tasks at this time https://twitter.com/...
  • @osanseviero Omar Sanseviero on x
    GPT 4, Claude, Alpaca, ChatGLM, PALM API... Can we pause for today? Too much to look at! 🙏
  • @wintonark Brett Winton on x
    A human typing flat-out for an hour will generate 3,000 or so tokens. At prevailing wages you will have paid $6 per 1k tokens. GPT-4 is priced at between $.03 and $.06 per 1k tokens (or $.06 to $.12 for 4x the context window.) 1/100th the price.
  • @danbarker Dan Barker on x
    I think these are the 2 changes in GPT-4 that will pick up most interest: 1) It can write 25,000 words of output - you could generate a unique, short novel in 10 seconds 2) It can pass the Uniform Bar Exam - the exam to determine readiness to practice law - with a top 10% score h…
  • @timnitgebru @timnitgebru on x
    Anyways, maybe accompany your GPT-4 hype train reading with this thread. https://twitter.com/...
  • @rsanti97 Santi Ruiz on x
    what explains GPT4's abysmal performance on AP English Literature? https://twitter.com/...
  • @karpathy Andrej Karpathy on x
    The GPT-4 developer livestream (https://www.youtube.com/...) was a great preview of new capability. Not sure I can think of a time where there was this much unexplored territory with this much new capability in the hands of this many users/developers. https://twitter.com/...
  • @mattturck Matt Turck on x
    Few understand how exhausting it is to provide Twitter thought leadership on epidiomology, oil, macroeconomics, war strategy, tokenomics, banking regulation just last weekend and now GPT-4, in rapid succession. Our heavy, heavy cross to bear.
  • @smokeawayyy @smokeawayyy on x
    GPT-4 came out on Tuesday instead of Thursday. 2 days is like 2 months at the current rate of progress. ...I need to adjust some timelines...
  • @thealexbanks Alex Banks on x
    1/ Visual inputs GPT-4 is a multimodal modal. It accepts both image and text inputs for text output. GPT-4's image recognition & understanding capabilities are mindblowing. https://twitter.com/...
  • @sama Sam Altman on x
    we are previewing visual input for GPT-4; we will need some time to mitigate the safety challenges.
  • @benmschmidt @benmschmidt on x
    Neural networks like GPT-4 are notoriously black boxes; the fact that their operations are unpredictable and inscrutable is one of *the* most important questions about whether and where they should be used. And now OpenAI is planting a standard to extend that mystery farther.
  • @lillysharples @lillysharples on x
    GPT-4 only scored a 3/45 on leetcode hard, we're safe for now https://twitter.com/...
  • @suhail @suhail on x
    We tried GPT-4 internally over the past few months and it is *a lot* better than meets the eye. I encourage folks to try it. World changing technology.
  • @billym2k Shibetoshi Nakamoto on x
    oh neat the new AI gpt-4 was dropped and it's many magnitudes more advanced than the previous one and also a us drone collided with a russian aircraft on international waters nothing potentially catastrophic or world ending here! 😅😅
  • @jon_barron Jon Barron on x
    From the GPT-4 paper. Not sure how I feel about this. https://twitter.com/...
  • @yascha_mounk Yascha Mounk on x
    It now seems quite clear that it's only a matter of time until artificial intelligence outperforms humans on skills we once thought of as distinctive to our species, such as composing a beautiful piece of music or writing a moving story. A major, melancholy milestone in history. …
  • @dellcam Dell Cameron on x
    i feel like OpenAI has somehow drastically underestimated the legal liability around this thing https://twitter.com/...
  • @0xfoobar @0xfoobar on x
    it costs $20/month for a superhuman AI assistant and you're worried about inflation
  • @abacaj Anton on x
    GPT-4 doing taxes with 32k seq length https://twitter.com/...
  • @_jasonwei Jason Wei on x
    IMO GPT-4 is a bigger leap than GPT-3 was. - GPT-3 advanced AI from task-specific models to a single prompted model that is task-general - GPT-4 is human-level on many hard tasks, and will signal a *societal* revolution where AI reaches every industry, starting with technology 🧵 …
  • @wintonark Brett Winton on x
    OpenAI announces GPT-4 here: https://openai.com/... Performance on human-benchmarks is rather remarkable. GPT-3.5 scored 10th percentile on the bar exam, GPT-4 hits the 90th percentile. On BC calculus it got the equivalent of a 4, good for college credit at 99% of colleges. https…
  • @refikanadol Refik Anadol on x
    Wow! Massive output changes for Dalle2. Looks like gpt4 will be a major game changer beyond just text. https://twitter.com/...
  • @nandodf @nandodf on x
    GPT-4 is absolutely brilliant. I love the focus on education and multilingual, which could empower many millions of disadvantaged people. Congratulations @gdb @ilyasut @sama @woj_zaremba and everyone at @OpenAI https://twitter.com/...
  • @jakebackpack Jacob Bacharach on x
    Incredibly funny to fail to note that you'd see similar pass/score rates among people if they were *allowed to use the internet to look up all the answers*! https://twitter.com/...
  • @wongmjane Jane Manchun Wong on x
    GPT-4 doesn't know it's GPT-4 https://twitter.com/...
  • @lopp Jameson Lopp on x
    GPT-4 released this week. Midjourney v5 released next week. The world will continue to get weirder!
  • @katecrawford Kate Crawford on x
    Crucial section from OpenAI's GPT4 paper: “AI systems will have even greater potential to reinforce entire ideologies, worldviews, truths and untruths, and to cement them or lock them in, foreclosing future contestation, reflection, and improvement.” https://cdn.openai.com/...
  • @tegmark Max Tegmark on x
    Although the Winograd test for AI is notoriously tougher than the Turing test, just-released GPT4 crushes it: https://twitter.com/...
  • @ammaar Ammaar Reshi on x
    GPT-4 is slowly rolling out to ChatGPT Plus subscribers today—just got access, it's capped at 100 messages every 4 hours. https://twitter.com/...
  • @tszzl Roon on x
    i find the creative capabilities of GPT4 astonishing, especially the long form coherence take a look at this one I whipped up. I asked for no more than a miltonian epic about a misaligned intelligence A Singularity Of Woe: An Epic In Twelve Cantos By GPT, After John Milton... htt…
  • @heydave7 Dave Lee on x
    GPT-4 is a big upgrade in language ability, commonsense reasoning, reading comprehension, coding tasks, exam testing, and more. https://twitter.com/...
  • @stevesi Steven Sinofsky on x
    @emollick Is this really that shocking? I think the most interesting scenarios are when there is not such an obvious corpus of highly structured training data that can be “matched” so readily. That is an innovation but is it really “something else”?
  • @coldhealing @coldhealing on x
    gpt-4 got a 2 on ap lit and ap lang but got a 4 on ap calc bc. it can't do leetcode yet but if i were a shape rotator i'd be worried https://twitter.com/...
  • @luxalptraum @luxalptraum on x
    This is more a condemnation of standardized tests than praise for GPT-4 https://twitter.com/...
  • @michael_nielsen Michael Nielsen on x
    Tidbits from the GPT-4 announcement post: https://openai.com/... It exceeds 80th percentile on quite a number of exams. Still bad at some - things like AP English lit https://twitter.com/...
  • @aeyakovenko @aeyakovenko on x
    GPT-4 > gpt-3.5 at generating bed time stories. Guess which one is 4 https://twitter.com/...
  • @emollick Ethan Mollick on x
    All from the GPT-4 whitepaper. https://openai.com/... https://twitter.com/...
  • @rasbt Sebastian Raschka on x
    Some takeaways after reading the GPT-4 paper: https://cdn.openai.com/... The good: 1) The model now accepts multimodal inputs: images and text 2) The paper is surprisingly honest: “While less capable than humans in many real-world scenarios...”, “GPT-4's capabilities and... https…
  • @autismcapital @autismcapital on x
    “Our mitigations have significantly improved many of GPT-4's safety properties compared to GPT-3.5. We've decreased the model's tendency to respond to requests for disallowed content by 82% compared to GPT-3.5.” Great, GPT-4 sucks 82% more than GPT-3. https://twitter.com/...
  • @manlikemishap Pamela Mishkin on x
    Evals are a key tool for us in tracking our progress on usefulness and safety. Submit an eval, get ahead on the GPT-4 waitlist! And reach out if you'd like AI to perform better on your language, context, or use-case and would like help constructing an eval. https://twitter.com/..…
  • @nonmayorpete Pete on x
    OpenAI just dropped GPT-4. Some bullets on the announcements: - Passes bar exam at top 10% of test takers - Passes many AP exams - Significantly outperforms leading models on benchmarks - Accepts text AND images
  • @saboo_shubham_ Shubham Saboo on x
    Data Analysis on the fly. Provide the image/text input or both and GPT-4 can answer all your questions just like an human analyst. https://twitter.com/...
  • @adamdangelo Adam D'Angelo on x
    GPT-4 is a major advance relative to ChatGPT and is the most powerful language model available to the world today. It is particularly strong at creative writing, problem solving, and instruction following. For example, GPT-4 solves this 2023 AIME (math contest) problem correctly.…
  • @simonw Simon Willison on x
    “gpt-4 has a context length of 8,192 tokens. We are also providing limited access to our 32,768-context (about 50 pages of text) version, gpt-4-32k” That's what I was most hoping for: enables things like summarization of full academic papers, rather than having to split them
  • @chrisalbon Chris Albon on x
    GPT-3.5 (ie chatgpt) vs GPT-4 (released today) https://twitter.com/...
  • @debarghya_das Deedy on x
    OpenAI's GPT-4 announcement is FINALLY HERE and its UNREAL! - Multimodal: Text + Image -> Text - SOTA on MMLU by a mile: 86.4 vs 75.2 for Flan-PALM - Able to score 1410/1600 on the SAT and 331/340 on the GMAT, beating 1260 and 331 on GPT-3.5 - Performance in 24 other languages...…
  • @lugaricano Luis Garicano on x
    Stunning performance on the most advanced exams. Better than 90% of aspiring lawyers on the Bar, better than 99% of aspiring graduate students on verbal intelligence! https://twitter.com/... https://twitter.com/...
  • @sjwhitmore Sam Whitmore on x
    excited to say that we've been able to build on top of gpt-4 for the past few months (one of the reasons i haven't shared many demos / screenshots). the model is *really* smart. thx @OpenAI ! twitter's gonna be crazy today but will share more soon :) — https://mydaemon.ai/ https:…
  • @drjimfan @drjimfan on x
    GPT-4 is HERE. Most important bits you need to know: - Multimodal: API accepts images as inputs to generate captions & analyses. - GPT-4 scores 90th percentile on BAR exam!!! And 99th percentile with vision on Biology Olympiad! Its reasoning capabilities are far more advanced... …
  • @emostaque Emad on x
    Great work by @OpenAI the higher context length (32k) in particular opens up a world of possibility even if not linear in cost understandably (reckon they'll fix that). Really appreciate how they took the time to thank everyone that contributed to the project too, team effort 👏🏾 …
  • @levie Aaron Levie on x
    GPT-4 is amazing 👏 https://twitter.com/...
  • @repligate Janus on x
    > We spent 6 months making GPT-4 safer and more aligned. GPT-4 is 82% less likely to respond to requests for disallowed content https://twitter.com/... https://twitter.com/...
  • @woj_zaremba Wojciech Zaremba on x
    I am extremely proud of what the hundreds of extremely talented, hardworking, caring people at OpenAI have built. Welcome GPT-4 to the world! We are happy to have you here. You will entertain us, bring us tears, make us laugh, and help us. Thank you. https://openai.com/...
  • @lugaricano Luis Garicano on x
    Here it is. GPT-4, which passes the Bar and College Calculus https://twitter.com/...
  • @tier10k @tier10k on x
    Introducing GPT-4, OpenAI's most advanced system https://t.co/ak6S0siai9
  • @kondrich2 Andrew Kondrich on x
    Excited to share what I've been working on at @openai! https://github.com/... is a framework for evaluating OpenAI models and an open-source eval registry. We will be granting GPT-4 access to those who submit high quality evals. Looking forward to your contributions!
  • @gdb Greg Brockman on x
    We're releasing GPT-4 — a large multimodal model (image & text in, text out) which is a significant advance in both capability and alignment. Still limited in many ways, but passes many qualification benchmarks like the bar exam & AP Calculus: https://openai.com/...
  • @gdb Greg Brockman on x
    GPT-4 as your personal tutor on @khanacademy. One of my personal dream applications (fun fact, I was at one point exploring starting a programming education company — always felt that more people would program if they had access to a great teacher). https://twitter.com/...
  • @gdb Greg Brockman on x
    GPT-4's vision capability for assisting blind & low vision users: https://twitter.com/...
  • @gdb Greg Brockman on x
    GPT-4 was, in many ways, our first whole-company project: https://openai.com/.... Jakub did a stellar job as lead for pretraining. https://twitter.com/...
  • @laurengoode Lauren Goode on x
    “While they've made a lot of progress, it's clearly not trustworthy,” says Oren Etzioni, prof emeritus at the UWash & the founding CEO of the Allen Institute for AI. “It's going to be a long time before you want any GPT to run your nuclear power plant.” https://www.wired.com/...
  • @brianstelter Brian Stelter on x
    GPT-4: “It's more accurate, but it still makes things up.” https://www.nytimes.com/...
  • @chafkin Max Chafkin on x
    in a bunch of these side by side comparisons gpt 3.5 seems the same or better than supposedly “advanced” system https://twitter.com/...
  • @chafkin Max Chafkin on x
    this explainer doesn't make much of a case that gpt-4 is better than gpt-3. and yet https://www.nytimes.com/... https://twitter.com/...
  • @john_bailey John Bailey on x
    GPT-4 with some impressive pass rates of a variety of exams. “We tested GPT-4 on a diverse set of benchmarks, including simulating exams that were originally designed for humans. We did no specific training for these exams.” https://cdn.openai.com/... https://twitter.com/...
  • @npew Peter Welinder on x
    GPT-4 can read images. Still in research preview, but the team is working hard on getting it ready for broader access. https://twitter.com/...
  • @whet Whet Moser on x
    dang, the thing where gpt 4 takes a picture of the inside of a fridge and suggests a meal is pretty wild https://www.nytimes.com/...
  • @shiraovide Shira Ovide on x
    Both of these AI jokes are bad. I feel ok about humans. https://www.nytimes.com/... https://twitter.com/...
  • @drjimfan @drjimfan on x
    Want early access to GPT-4? Do it now: https://github.com/... is an official framework for evaluating OpenAI models. They will grant GPT-4 access to those who submit high quality evals. Thanks to my friend Andrew Kondrich @kondrich2 who built this initiative at OpenAI!
  • @e0m Evan Morikawa on x
    If you see something GPT-4 can't do well, or think you can prove a fundamental deficiency, contribute evals! This is by far the best way to help close these skill gaps. Internally we use evals to guide enormous amounts of model development. https://github.com/...
  • @swyx @swyx on x
    as LLMs grow and grow and grow in capabilities, it is getting more impt to have good model evaluation/benchmarking frameworks. OpenAI is also releasing their eval framework, fully MIT licensed: https://github.com/... Used by Stripe and well documented. Runs MMLU in 189 LOC https:…
  • @officiallogank @officiallogank on x
    We are giving priority GPT-4 access to those who contribute evals to our new evals repo: https://github.com/... Here, you can write tests for the model so we can improve things over time.