1. X
  2. Unstructured
Log inSign up
Unstructured
1,621 posts
Unstructured profile banner
user avatar

Unstructured

@UnstructuredIO
Stop dilly-dallying. Get your data. 👉🏼 Get Started: unstructured.io
San Francisco, CA
unstructured.io
Joined August 2022
152
Following
6,398
Followers
RepliesRepliesMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • user avatar
    Unstructured
    @UnstructuredIO
    21h
    Technical docs are full of code, and code isn't just another paragraph of text. Unstructured treats code as its own element type, CodeSnippet. When we parse docs, textbooks, or API references, code gets detected, extracted with its formatting intact, and stored separately from
    Image
  • user avatar
    Unstructured
    @UnstructuredIO
    Aug 27
    Move a doc out of @IBM FileNet and its permissions usually get left behind. We don't let that happen 😎 Unstructured's FileNet connector captures the full ACL (read/update/delete, allow/deny, up to 1,000 entries) at ingestion and attaches it to every element as
    Image
  • user avatar
    Unstructured
    @UnstructuredIO
    Aug 26
    Why do tables always fall apart when you pull them out of a document? Most parsers flatten everything to plain text, and the structure goes with it. We don't. Unstructured keeps tables as structured HTML in the text_as_html field, so every row, column, and header relationship
    Image
  • user avatar
    Unstructured
    @UnstructuredIO
    Aug 25
    Ever tried pulling the formulas out of a century-old physics paper? 😵‍💫 Formulas are some of the hardest content to parse accurately. Our High Res partitioner runs object detection to find the formula regions, then hands those bounding boxes to our VLM-based enrichments for
    Image
  • user avatar
    Unstructured
    @UnstructuredIO
    Aug 24
    Unstructured doesn't just hand you a wall of text. It hands you elements: Titles, NarrativeText, ListItems, Tables, and more, each tagged with metadata like page number, coordinates, and languages. That structure is what makes everything downstream possible: filter by element
    Image
Advertisement
Advertisement