Give AlbumentationsX a star on GitHub — it powers this leaderboard

Star on GitHub

Unstructured-IO/unstructured

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

Stars: 14,094Language: HTML
Unstructured-IO/unstructured - GitHub Repository | PyPI Leaderboard