Inside the ChatGPT Leak: How ChatGPT Ranks Chunks From Fetched Pages [CHUNKING PROOF]
Written by Metehan Yeşilyurt. https://www.linkedin.com/pulse/inside-chatgpt-leak-how-ranks-chunks-from-fetched-pages-ye%C5%9Filyurt--qrsyfLast week, we came across incredible findings in the SSEs. Scores, schema, and more. I shared the initial details with you two days ago.
Special thanks to our great research team at Peec AI .
Now, we will continue to share further details in the form of an article.
The ChatGPT SSE dump reveals how web content is retrieved: pages fetched during search are segmented into discrete passages and individual chunks are scored with neural relevance weights.
How we got here
With the SSE dump we can see the retrieval data behind any normal ChatGPT prompt. Shopping prompts were the open question, so we tried one. We started with a single comparison prompt: "What are the best wireless noise cancelling headphones under $300 right now? Compare the Sony WH-1000XM5, Bose QuietComfort Headphones and Sennheiser Momentum 4." ChatGPT answered with 5 product cards (Sony WH-1000XM5, Bose QuietComfort Headphones, Sennheiser Momentum 4 Wireless, Sony ULT Wear, Soundcore Space One Pro). Behind those cards the same turn ran 4 web searches and fetched 96 pages. We pulled the web retrieval results alongside the product cards, and that is where the SoundGuys record comes from.
-IMPORTANT- Everything can change tomorrow or next month. All of our findings are coming from the ChatGPT UI, not API.
As far as we can tell, ChatGPT's entire retrieval system operates in a multi-layered, hybrid. A/B variants from Statsig and even model selection alongside "Instant," "Medium," and "High" settings change the retrieval configuration. In other words, it is a dynamic system.
The Finding in Brief
There has been intense debate about whether ChatGPT ranks whole web pages or chunks. Many claimed that ChatGPT evaluates pages only as cohesive units. Others claimed something different. Chunking "was" one of the biggest black boxes in our industry.
Now we have the proof.
The SSE dump gives us the conclusive answer: both claims are incomplete. ChatGPT uses a multi-stage hybrid retrieval hierarchy. If anyone mentioned "multi-stage" hybrid retrieval before, that was absolutely correct. Period.
The Core Mechanism
Based on the SSE dump, ChatGPT pulls pre-scored passage chunks from its search index cache. When a page is already indexed, the passage chunks and their neural relevance scores are carried straight into the fetch payload and normalized to 1.0000. Fresh or unindexed pages are split into hundreds of small fragments at runtime.
Read more about ChatGPT Index: https://peec.ai/blog/chatgpt-built-its-own-search-index
The SoundGuys Evidence
For the comparison prompt above, one of the 4 web searches was best noise cancelling headphones under $300 2026 Sony WH-1000XM5 Bose QuietComfort Momentum 4, restricted to rtings.com, tomsguide.com and soundguys.com. It returned this SoundGuys URL:
https://www.soundguys.com/sony-wh-1000xm5-vs-bose-noise-cancelling-headphones-700-71868/The search result record holds 3 passages in snippet::parts and 3 scores in snippet::scores. The fetched_pages record for the same URL holds the same 3 passages, in the same order, with rescaled scores.
The Mathematical Proof: Max Score Normalization
The third passage had the highest raw score from the neural ranker sonic-re-fh-chunk-ev3-sx4: 0.9983865023.
In the page_contents structure that same passage became exactly 1.0000. The other two were rescaled by roughly the same factor:
How sure are we?
The turn fetched 96 pages. 76 of them are pre-indexed pages(labrador) whose search result record also carries raw snippet scores, so the raw vs page comparison is only possible for those 76 (the other 20 are fresh fetches without raw scores). We checked all 76. Chunk count and order match on every page. The top chunk is exactly 1.0 for web, PDF, Reddit, YouTube and Wikipedia results (68 pages) and exactly 0.5 for news results (8 pages). Ratios between chunks hold to about four decimals (median deviation 0.00001, worst 0.0008). The multiplier is not bit-exact: it drifts in the sixth decimal from chunk to chunk, so the exact arithmetic is not visible in the SSE dump. We're still working on it.
The Code Area: Inside the SSE Dump
This is the raw search result record for one SoundGuys URL, exactly as stored. snippet::parts holds the page passages, snippet::scores holds one score per passage, in the same order. The first passage is shown in full; the other two are shortened only for display.
Citation Markup Resolution
Why is parts[0] 1,289 characters in the search result but 1,279 characters in the fetched page? The text is the same. The only difference is the inline citation anchors.
In the search result the passage carries anchors like 【28†Image: Bose Noise Canceling Headphones 700-6-2】 and 【29†Our standard battery test】. In the page representation handed to the model, the anchor wrappers are removed and the link text stays.
-IMPORTANT NOTE AGAIN- Everything can change tomorrow or next month. All of our findings are coming from the ChatGPT UI, not API.
As far as we can tell, ChatGPT's entire retrieval system operates in a multi-layered, hybrid. A/B variants from Statsig and even model selection alongside "Instant," "Medium," and "High" settings, alter the retrieval configuration. In other words, it is a dynamic system.
We're still working on the full dump. Stay tuned.
FAQs
-Why don't you share everything exactly as it is? Are you holding anything back? No, it might look that way right now, but a single JSON dump can exceed 200,000 lines(for only 1 prompt, we're using different settings like instant, medium, high, shopping, comparison queries, etc). We can generate these dumps for every prompt we send via the UI. Reviewing them really takes time. And we didn't automate the process because we want to keep testing accounts safe. So, every a/b gate opens new doors.
-What do you have exactly? It looks like a great stuff. Tons of work. Web retrieval, shopping (Google shopping scraping + OAI Shopping) dumps.
-Why do you review them manually? Doesn't summarizing them with Claude, Gemini, or ChatGPT work? No, it doesn't. Due to context window limitations, parts inevitably get skipped; the old-fashioned way. Sitting down as a team to read and analyze them is better.
-Did you start running experiments immediately based on your findings? Certainly; we have actively begun conducting experiments based on the initial findings, so we recommend following Peec AI.
-What are the three keywords that stand out amidst this entire leak? Semantics, originality (but truly original), and authority.
Make sure you follow David Konitzny , Jan Ehrlinspiel , Malte Landwehr , Metehan Yeşilyurt , Tomek Rudzki
Get new research on AI search, SEO experiments, and LLM visibility delivered to your inbox.
Powered by Substack · No spam · Unsubscribe anytime