AI / GEO
เครื่องมือวิเคราะห์ Chunk Extractability ของเนื้อหา
AI answer engine และระบบ RAG ไม่ได้อ่านทั้งหน้าเว็บของคุณ แต่จะดึงเนื้อหาทีละ chunk แล้วแสดงแยกเดี่ยว ๆ เครื่องมือนี้แบ่งเนื้อหาของคุณด้วยวิธีเดียวกัน และให้คะแนนแต่ละ chunk ว่ายังเข้าใจได้หรือไม่เมื่อไม่มีบริบทรอบข้าง
หลักการเหล่านี้ปรับให้เหมาะกับข้อความภาษาอังกฤษ การนับคำและการตรวจโครงสร้างยังใช้ได้กับภาษาที่เขียนด้วยอักษรละตินอื่น ๆ แต่การตรวจคำสรรพนามและวลีเป็นแบบเฉพาะภาษาอังกฤษ
เครื่องมือนี้ตรวจอะไรบ้าง
AI Overviews, ChatGPT, Perplexity และการค้นหาแบบ RAG ล้วนทำงานคล้ายกัน คือแบ่งหน้าเว็บออกเป็นชิ้นส่วน (chunk) เก็บแยกกัน แล้วดึงมาแสดงเพียงหนึ่งหรือสอง chunk ที่ตอบคำถามได้ ส่วนที่เหลือของหน้าจะไม่ถูกส่งให้โมเดลเลย chunk ที่ต้องพึ่งพาสิ่งที่กล่าวไว้ก่อนหน้าในหน้าเว็บ เริ่มต้นแบบขาดตอนกลางความคิด หรืออาศัยตัวย่อที่ไม่ได้อธิบาย จะทำให้สับสนหรือถูกตัดทิ้งในขั้นตอนการดึงข้อมูลนี้ แม้ว่าทั้งหน้าเว็บจะอ่านได้ราบรื่นสมบูรณ์ก็ตาม เครื่องมือนี้จำลองขั้นตอนการแบ่ง chunk นั้นขึ้นมาใหม่ และให้คะแนนแต่ละ chunk ที่ได้ในแง่ความสมบูรณ์ในตัวเอง
How chunks are built and scored
Nothing is guessed from a template — every chunk comes from your own pasted text, split the way a heading-aware ingestion pipeline would split it (roughly 150-240 words per chunk, breaking at headings first).
Heading-anchored chunking
Paragraphs are grouped under their nearest heading into passages of roughly 150-240 words each. A single paragraph longer than that is flagged on its own, since an automatic chunker has no natural break point inside it.
Self-containment checks
Each chunk's opening sentence is checked for a dangling pronoun ("This", "It", "They"...) with no antecedent inside that same chunk, and its full text is scanned for backward-looking phrases like "as mentioned above" that only make sense next to a chunk the reader will never see.
Cross-chunk acronym tracking
Every acronym or abbreviation is tracked across the whole document. If it is spelled out in one chunk but reused unexplained in another, that second chunk is flagged — a retrieval system will show it alone, without the definition.
Length and mid-thought cutoffs
Chunks under 40 words are usually too thin to carry a complete idea; chunks that end on a colon or a conjunction ("...and", "...because") read as cut off mid-sentence to anyone shown only that passage.
คำถามที่พบบ่อย
เหมือนกับเครื่องมือตรวจความอ่านง่ายหรือความหนาแน่นคำสำคัญหรือไม่
ไม่เหมือน ความอ่านง่ายวัดว่ามนุษย์อ่านข้อความได้สบายหรือไม่ ส่วนเครื่องมือนี้วัดว่า chunk หนึ่งชิ้น เมื่อถูกแสดงเดี่ยว ๆ โดยไม่มีหน้าเว็บล้อมรอบ ยังเข้าใจได้หรือไม่ ซึ่งเป็นปัญหาคนละแบบที่สำคัญเพราะ AI answer engine ดึงและแสดง chunk ที่แยกออกมาต่างหาก
What counts as a "citation-ready" chunk?
A chunk scores 100 minus penalties for each issue found (a severe issue like a dangling pronoun costs more than a minor one like running long). It is marked citation-ready once it scores 80 or higher with no severe issues.
My content reads fine to a human — why are chunks flagged?
A full page reads fine because a human keeps context across paragraphs automatically. A retrieval system does not: it stores and shows one chunk in isolation, so anything that depends on an earlier sentence, an earlier chunk, or an earlier definition can fail silently even though the page itself is well written.
Does this tool send my text anywhere?
Your pasted text is sent once, over HTTPS, to our server to run the chunking and scoring (needed to keep the logic server-side and consistent), and is not stored beyond the request.