/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

DeepSeek releases DeepSeek-OCR, a vision language model designed for efficient vision-text compression, enabling longer contexts with less compute

the new frontier of OCR from @deepseek_ai , exploring optical context compression for LLMs, is running blazingly fast on vLLM ⚡ (~2500 tokens/s on A100-40G) — powered by vllm==0.8.5 for day-0 model support.  🧠 Compresses visual contexts up to 20× while keeping 97% OCR accuracy at <10×... Adina Yakup / @adinayakup : DeepSeek-OCR is out 🔥 https://huggingface.co/... ✨High-accuracy OCR - MIT license ✨Fast GPU inference (FlashAttention 2, BF16) ✨Docs > Markdown ✨Works with transformers @teortaxestex : I failed to parse the ambition of this release. DeepSeek Contexts Optical Compression is not just a good fast OCR, not just «we want to train V4/V5 on all Anna and DuXiu». It's exactly what it says in the title. And more. For starters, think of realtime computer-use agents. [image] Shai Wininger / @shai_wininger : Interesting. We've done the same internally a while back after noticing better context extraction precision with pixels rather than vector shaped post script text and ASCII. @teortaxestex : Massively unexpected update from DeepSeek: a powerful, high-compression MoE OCR model. > In production, DeepSeek-OCR can generate 33 million pages of data per day for LLMs/VLMs using 20 nodes (x8 A100-40G). They want ALL the tokens. You're welcome to have some too. [image] Aamir Shakir / @aaxsh18 : the deepseek ocr paper is interesting since it gives us a glimpse of the future. we will treat data in their native format instead of going back to text. but like all recent open source OCR models, it is not good on properly parsing complex documents. Brian Roemmele / @brianroemmele : BOOOOOOOM! CHINA DEEPSEEK DOES IT AGAIN! An entire encyclopedia compressed into a single, high-resolution image! — A mind-blowing breakthrough. DeepSeek-OCR, unleashed an electrifying 3-billion-parameter vision-language model that obliterates the boundaries between text and [image] @godofprompt : 🚨 DeepSeek just did something wild.  They built an OCR system that compresses long text into vision tokens literally turning paragraphs into pixels.  Their model, DeepSeek-OCR, achieves 97% decoding precision at 10× compression and still manages 60% accuracy even at 20×.  That means one image can represent entire documents using a fraction of the tokens an LLM would need... Alexander Doria / @dorialexander : So longer read of DeepSeek-OCR It's an engineering achievement. It has been suspected for a while that VLM/OCR models could be significantly smaller. The pre-VLM state of the art, Google Cloud OCR would not be more than a 100m model. More recently, relatively small open weights Ray Fernando / @rayfernando1337 : This is the JPEG moment for AI.  Optical compression doesn't just make context cheaper.  It makes AI memory architectures viable.  Training data bottlenecks?  Solved.  - 200k pages/day on ONE GPU - 33M pages/day on 20 nodes - Every multimodal model is data-constrained.  Not anymore... @maziyarpanahi : DeepSeek-OCR on doctor's hand written note! [image] Simon Willison / @simonw : DeepSeek released a new OCR model - I got it working on an NVIDIA Spark (CUDA + ARM64) by starting a Docker container, running Claude Code as root with “claude —dangerously-skip-permissions” and telling it to figure it out Took 4 prompts and 40 minutes https://simonwillison.net/... LinkedIn: Salome Khvistani : 📑 The newly released DeepSeek-OCR paper explores Optical Context Compression as a method to represent textual information through the visual modality. … Bluesky: Tim Kellogg / @timkellogg.me : i think this is the crux of DeepSeek-OCR  —  1. (text) context gets longer as you add words  —  2. long context is quadratic  —  3. you can fit lots of words in an image  —  4. if you use encoder-decoder architecture, your tokens encode a ton of information [embedded post] Tim Kellogg / @timkellogg.me : DeepSeek-OCR  —  a tiny 3B-A0.5B MoE OCR model that runs fast on a single A100 40GB with very high precision and excellent compression  —  why it's cool — they use images as a way to compress text and get around the O(n^2)  —  huggingface.co/deepseek-ai/...  [image] Forums: Hacker News : DeepSeek OCR r/LocalLLaMA : The Innovations in DeepSeek OCR r/LocalLLaMA : DeepSeek releases DeepSeek OCR

The Decoder Jonathan Kemper

Context & Ripple Effects

DeepSeek-OCR extends DeepSeek’s earlier pattern of releasing models under permissive terms, including the MIT-licensed V3-0324 release. Here, the emphasis shifts from model-scale competition to reducing the cost of carrying document and visual information into language-model workflows.

The release also provides the first reference point for the later OCR 2 upgrade, suggesting optical context compression is becoming a continuing product and research line rather than a one-off OCR utility.

First-order effects

  • Developers can access and integrate an MIT-licensed, Hugging Face-hosted OCR model through Transformers, with vLLM support aimed at rapid GPU inference.
  • For document-heavy prompts, representing text visually can reduce the context burden: DeepSeek reports up to 20× visual-context compression while retaining 97% OCR accuracy at under 10× compression.

Second-order effects

  • OCR and document-AI providers face a more explicit efficiency benchmark: quality must increasingly be weighed against the compute and context cost of passing extracted information to an LLM.
  • Teams processing large document collections can evaluate visual compression as an alternative to sending full text contexts, potentially shifting demand toward inference stacks optimized for this workflow.

Third-order effects

  • If accuracy holds across production documents, context management may move beyond tokenization alone: modality choice becomes another lever for controlling inference cost and usable context length.
  • Permissively licensed efficiency techniques can concentrate differentiation in deployment, data pipelines, and serving performance rather than model access alone.

The trend: AI inference is increasingly being optimized around the cost of context, with model and serving designs seeking to fit more usable information into a fixed compute budget.

Discussion

  • @doodlestein Jeffrey Emanuel on x
    DeepSeek just released a pretty shocking new paper. They really buried the lede here by referring to it simply as DeepSeek OCR. While it's a very strong OCR model, the purpose of it and the implications of their approach go far beyond what you'd expect of “yet another OCR
  • @karpathy Andrej Karpathy on x
    I quite like the new DeepSeek-OCR paper.  It's a good OCR model (maybe a bit worse than dots), and yes data collection etc., but anyway it doesn't matter.  The more interesting part for me (esp as a computer vision at heart who is temporarily masquerading as a natural language pe…
  • @dbreunig Drew Breunig on x
    This is such an interesting idea. The first step of building or using LLMs has always been making text cosplay as pixel-problems (gotta fit words into a component that evolved to render videogames). Why not skip the encoding and just put the text in an image? 🫠
  • @eliebakouch Elie on x
    DeepSeek-OCR has some weird architectural choices for the LLM decoder: DeepSeek3B-MoE-A570M -> uses MHA, no MLA (not even GQA?) -> 2 shared experts (like DeepSeek V2, but V3 only has 1) -> quite low sparsity, activation ratio is 12.5%. For V3 it's 3.52%, for V2 it's 5% -> not [im…
  • @_tobiaslee Lei Li on x
    DeepSeek-OCR: Exploring the boundaries of visual-text compression. Ambitious! They might use 10X (near-lossless) compressed vision tokens to replace the KV cache of dialog histories. https://github.com/... [image]
  • @reach_vb @reach_vb on x
    pretty cool to see that the model still retains its visual understanding capabilities 🤯 [image]
  • @inductionheads @inductionheads on x
    Thinking visually will help these models make better PowerPoint slides (the final frontier)
  • @thingshiddenn Moriarty on x
    Months ago, I commented (forecast) that one of DeepSeek next breakthroughs would be cracking OCR, creating a cutting edge approach to it. OCR is a way for machines to literally read text from documents, ESPECIALLY PDFS. DeepSeek comprensses this to insane eficiency.
  • @1littlecoder @1littlecoder on x
    Deepseek just released a new paper! The most interesting aspect is treating OCR as Optical Compression! Instead of storing or processing every text token directly (as LLMs do, which scales quadratically with length), DeepSeek-OCR represents text visually, for example, a page [ima…
  • @jenzhuscott Jen Zhu on x
    Is it just me or DeepSeek OCR paper feels like something they've done a while ago but only decided to release now? What are they working on?? [image]
  • @cjzafir CJ Zafir on x
    Deepseek did it again!!! Transformed context, long form memory, and RAG with their new “optical compression” technique. If you guys are not realizing, open source AI is bringing 99% of the innovation to AI tech. While closed source AI is hyping up the tech, raising billions in
  • @harveenchadha Harveen Singh Chadha on x
    deepseek released a new OCR model demonstrating a way to compress images into a smaller set of vision tokens 10× smaller while still achieving 97% accuracy, even at 20× compression retains around 60% accuracy can generate 200k+ pages/day on a single A100-40G !! [image]
  • @adithya_s_k Adithya S K on x
    Fun fact deepseek ocr is built by the same team that released GOT OCR a while back
  • @dorialexander Alexander Doria on x
    We still have to run the internal bench but not surprised. Complex parsing is primarily a data/synth recipes problem.
  • @untitled01ipynb @untitled01ipynb on x
    copyright? where we're going there's no such thing [image]
  • @chatgpt21 Chris on x
    DeepSeek-OCR, a 3B-param vision-language model, just dropped! Specializing in OCR & markdown conversion, it's open source (MIT) on Hugging Face https://huggingface.co/...
  • @tokenbender @tokenbender on x
    what a bold direction by deepseek once again. they took “a picture is worth a thousand words” literally or the idea of “photographic memory” if i am to commit the crime of anthropomorphisation. [image]
  • @casper_hansen_ Casper Hansen on x
    NEW DeepSeek OCR model that outperforms dots ocr while prefilling 3x less tokens [image]
  • @brianroemmele Brian Roemmele on x
    Government supplied FREE training data, because that government actually has a Manhattan Project to preserve non-internet data is propelling this country's AI rapidly. You hear me up here doing all I can to alert the US. Testing, testing, is this thing on? 🚨‼️
  • @jiqizhixin @jiqizhixin on x
    Compress everything visually! DeepSeek has just released DeepSeek-OCR, a state-of-the-art OCR model with 3B parameters. Core idea: explore long-context compression via 2D optical mapping. Architecture: - DeepEncoder → compresses high-res inputs into few vision tokens; - [image]
  • @rohanpaul_ai Rohan Paul on x
    DeepSeek-OCR just dropped. 🔥 Sets a new standard for open-source OCR A 3B-parameter vision-language model designed for high-performance optical character recognition and structured document conversion. - Can parse and re-render charts in HTML - Optical Context Compression: [image…
  • @dr_singularity Dr Singularity on x
    Big AI progress “DeepSeek figured out how to get 10x better compression using vision tokens than with text tokens. So you could theoretically store those 10k words in just 1,500 of their special compressed visual tokens.”
  • @reach_vb @reach_vb on x
    Letsss gooo! DeepSeek just released a 3B OCR model on Hugging Face 🔥 Optimised to be token efficient AND scale ~200K+ pages/day on A100-40G Same arch as DeepSeek VL2 Use it with Transformers, vLLM and more 🤗 https://huggingface.co/...
  • @mervenoyann Merve on x
    DeepSeek-OCR is out! 🔥 my take ⤵️ > pretty insane it can parse and re-render charts in HTML > it uses CLIP and SAM features concatenated, so better grounding > very efficient per vision tokens/performance ratio > covers 100 languages [image]
  • @iscienceluvr Tanishq Mathew Abraham, Ph.D. on x
    DeepSeek released an OCR model today. Their motivation is really interesting: they want to use visual modality as an efficient compression medium for textual information, and use this to solve long-context challenges in LLMs. Of course, they are using it to get more training [ima…
  • @brianroemmele Brian Roemmele on x
    Now I know why a few companies in China wanted to hire me with a “name your offer” situation including a house, cars, a lab, staff and “work on what you want”.
  • @vllm_project @vllm_project on x
    🚀 DeepSeek-OCR — the new frontier of OCR from @deepseek_ai , exploring optical context compression for LLMs, is running blazingly fast on vLLM ⚡ (~2500 tokens/s on A100-40G) — powered by vllm==0.8.5 for day-0 model support.  🧠 Compresses visual contexts up to 20× while keeping 97…
  • @adinayakup Adina Yakup on x
    DeepSeek-OCR is out 🔥 https://huggingface.co/... ✨High-accuracy OCR - MIT license ✨Fast GPU inference (FlashAttention 2, BF16) ✨Docs > Markdown ✨Works with transformers
  • @teortaxestex @teortaxestex on x
    I failed to parse the ambition of this release. DeepSeek Contexts Optical Compression is not just a good fast OCR, not just «we want to train V4/V5 on all Anna and DuXiu». It's exactly what it says in the title. And more. For starters, think of realtime computer-use agents. [imag…
  • @shai_wininger Shai Wininger on x
    Interesting. We've done the same internally a while back after noticing better context extraction precision with pixels rather than vector shaped post script text and ASCII.
  • @teortaxestex @teortaxestex on x
    Massively unexpected update from DeepSeek: a powerful, high-compression MoE OCR model. > In production, DeepSeek-OCR can generate 33 million pages of data per day for LLMs/VLMs using 20 nodes (x8 A100-40G). They want ALL the tokens. You're welcome to have some too. [image]
  • @aaxsh18 Aamir Shakir on x
    the deepseek ocr paper is interesting since it gives us a glimpse of the future. we will treat data in their native format instead of going back to text. but like all recent open source OCR models, it is not good on properly parsing complex documents.
  • @brianroemmele Brian Roemmele on x
    BOOOOOOOM! CHINA DEEPSEEK DOES IT AGAIN! An entire encyclopedia compressed into a single, high-resolution image! — A mind-blowing breakthrough. DeepSeek-OCR, unleashed an electrifying 3-billion-parameter vision-language model that obliterates the boundaries between text and [imag…
  • @godofprompt @godofprompt on x
    🚨 DeepSeek just did something wild.  They built an OCR system that compresses long text into vision tokens literally turning paragraphs into pixels.  Their model, DeepSeek-OCR, achieves 97% decoding precision at 10× compression and still manages 60% accuracy even at 20×.  That me…
  • @dorialexander Alexander Doria on x
    So longer read of DeepSeek-OCR It's an engineering achievement. It has been suspected for a while that VLM/OCR models could be significantly smaller. The pre-VLM state of the art, Google Cloud OCR would not be more than a 100m model. More recently, relatively small open weights
  • @rayfernando1337 Ray Fernando on x
    This is the JPEG moment for AI.  Optical compression doesn't just make context cheaper.  It makes AI memory architectures viable.  Training data bottlenecks?  Solved.  - 200k pages/day on ONE GPU - 33M pages/day on 20 nodes - Every multimodal model is data-constrained.  Not anymo…
  • @maziyarpanahi @maziyarpanahi on x
    DeepSeek-OCR on doctor's hand written note! [image]
  • @simonw Simon Willison on x
    DeepSeek released a new OCR model - I got it working on an NVIDIA Spark (CUDA + ARM64) by starting a Docker container, running Claude Code as root with “claude —dangerously-skip-permissions” and telling it to figure it out Took 4 prompts and 40 minutes https://simonwillison.net/.…
  • @timkellogg.me Tim Kellogg on bluesky
    i think this is the crux of DeepSeek-OCR  —  1. (text) context gets longer as you add words  —  2. long context is quadratic  —  3. you can fit lots of words in an image  —  4. if you use encoder-decoder architecture, your tokens encode a ton of information [embedded post]
  • @timkellogg.me Tim Kellogg on bluesky
    DeepSeek-OCR  —  a tiny 3B-A0.5B MoE OCR model that runs fast on a single A100 40GB with very high precision and excellent compression  —  why it's cool — they use images as a way to compress text and get around the O(n^2)  —  huggingface.co/deepseek-ai/...  [image]
  • r/LocalLLaMA r on reddit
    The Innovations in DeepSeek OCR
  • r/LocalLLaMA r on reddit
    DeepSeek releases DeepSeek OCR