Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

chore(deps): bump unstructured from 0.10.27 to 0.16.16 #280

Closed
wants to merge 1 commit into from

Conversation

dependabot[bot]
Copy link
Contributor

@dependabot dependabot bot commented on behalf of github Jan 28, 2025

Bumps unstructured from 0.10.27 to 0.16.16.

Release notes

Sourced from unstructured's releases.

0.16.16

Enhancements

Features

  • Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.

Fixes

  • Add auto-download for NLTK for Python Enviroment When user import tokenize, It will automatic download nltk data from tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.
  • Correctly patch pdfminer to avoid PDF repair. The patch applied to pdfminer's parser caused it to occasionally split tokens in content streams, throwing PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.
  • Drop usage of ndjson dependency

0.16.15

  • Update unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libs
  • Update pdfminer-six to 20240706

0.16.14

Enhancements

Features

Fixes

  • Fix an issue with multiple values for infer_table_structure when paritioning email with image attachements the kwarg calls into partition to partition the image already contains infer_table_structure. Now partition function checks if the kwarg has infer_table_structure already

0.16.13

Enhancements

  • Add character-level filtering for tesseract output. It is controllable via TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.

Features

Fixes

  • Fix NLTK Download to use nltk assets in docker image
  • removed the ability to automatically download nltk package if missing

0.16.12

Enhancements

  • Prepare auto-partitioning for pluggable partitioners. Move toward a uniform partitioner call signature so a custom or override partitioner can be registered without code changes.
  • Add NDJSON file type support.

Features

Fixes

  • Base image has been updated.
  • Upgrade ruff to latest. Previously the ruff version was pinned to <0.5. Remove that pin and fix the handful of lint items that resulted.
  • CSV with asserted XLS content-type is correctly identified as CSV. Resolves a bug where a CSV file with an asserted content-type of application/vnd.ms-excel was incorrectly identified as an XLS file.

... (truncated)

Changelog

Sourced from unstructured's changelog.

0.16.16

Enhancements

Features

  • Vectorize layout (inferred, extracted, and OCR) data structure Using np.ndarray to store a group of layout elements or text regions instead of using a list of objects. This improves the memory efficiency and compute speed around layout merging and deduplication.

Fixes

  • Add auto-download for NLTK for Python Enviroment When user import tokenize, It will automatic download nltk data from tokenize.py file. Added AUTO_DOWNLOAD_NLTK flag in tokenize.py to download NLTK_DATA.
  • Correctly patch pdfminer to avoid PDF repair. The patch applied to pdfminer's parser caused it to occasionally split tokens in content streams, throwing PDFSyntaxError. Repairing these PDFs sometimes failed (since they were not actually invalid) resulting in unnecessary OCR fallback.
  • Drop usage of ndjson dependency

0.16.15

Enhancements

Features

Fixes

  • Update unstructured-inference to 0.8.6 in requirements which removed layoutparser dependency libs
  • Update pdfminer-six to 20240706

0.16.14

Enhancements

Features

Fixes

  • Fix an issue with multiple values for infer_table_structure when paritioning email with image attachements the kwarg calls into partition to partition the image already contains infer_table_structure. Now partition function checks if the kwarg has infer_table_structure already

0.16.13

Enhancements

  • Add character-level filtering for tesseract output. It is controllable via TESSERACT_CHARACTER_CONFIDENCE_THRESHOLD environment variable.

Features

Fixes

  • Fix NLTK Download to use nltk assets in docker image
  • removed the ability to automatically download nltk package if missing

0.16.12

Enhancements

  • Prepare auto-partitioning for pluggable partitioners. Move toward a uniform partitioner call signature so a custom or override partitioner can be registered without code changes.
  • Add NDJSON file type support.

Features

... (truncated)

Commits

Dependabot compatibility score

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot merge will merge this PR after your CI passes on it
  • @dependabot squash and merge will squash and merge this PR after your CI passes on it
  • @dependabot cancel merge will cancel a previously requested merge and block automerging
  • @dependabot reopen will reopen this PR if it is closed
  • @dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)

@dependabot dependabot bot added the chore label Jan 28, 2025
@github-actions github-actions bot added the dependencies Pull requests that update a dependency file label Jan 28, 2025
@dependabot dependabot bot force-pushed the dependabot/pip/unstructured-0.16.16 branch 2 times, most recently from 6f26855 to 687a91a Compare January 29, 2025 04:38
Bumps [unstructured](https://github.com/Unstructured-IO/unstructured) from 0.10.27 to 0.16.16.
- [Release notes](https://github.com/Unstructured-IO/unstructured/releases)
- [Changelog](https://github.com/Unstructured-IO/unstructured/blob/main/CHANGELOG.md)
- [Commits](Unstructured-IO/unstructured@0.10.27...0.16.16)

---
updated-dependencies:
- dependency-name: unstructured
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <[email protected]>
@dependabot dependabot bot force-pushed the dependabot/pip/unstructured-0.16.16 branch from 687a91a to 33edf5a Compare January 29, 2025 05:13
Copy link
Contributor Author

dependabot bot commented on behalf of github Jan 29, 2025

Superseded by #294.

@dependabot dependabot bot closed this Jan 29, 2025
@dependabot dependabot bot deleted the dependabot/pip/unstructured-0.16.16 branch January 29, 2025 16:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
chore dependencies Pull requests that update a dependency file
Projects
None yet
Development

Successfully merging this pull request may close these issues.

0 participants