Linux workstation

Upstream software project

python-tokenizers

Provides an implementation of today's most used tokenizers

About python-tokenizers

Provides an implementation of today's most used tokenizers

This project links 5 native package records across 4 recorded operating-system releases. Compare the retained versions and architectures below, then open the package for your own release.

These are catalog observations, not a guarantee of installation, compatibility, or upstream support.

Project pictures and package coverage

python-tokenizers repository preview from GitHub
Project repository preview, supplied by GitHub. Open source repository
Fedora 43: 1 package records; Fedora 44: 1 package records; openSUSE Leap 16.0: 1 package records; openSUSE Tumbleweed: 2 package records. Catalog coverage diagram, not an application screenshot.python-tokenizers: recorded package coverageFedora 431 recordsFedora 441 recordsopenSUSE Leap 16.01 recordsopenSUSE Tumbleweed2 records
OpenFactory diagram of linked package records. It is not an application screenshot.

Project identity

Project
python-tokenizers
Publisher
Not authoritatively mapped
Native package records
5
Operating systems
fedora-43, fedora-44, opensuse-leap-16-0, opensuse-tumbleweed
License expression
Apache-2.0
Metadata completeness
100/100 (not a software quality rating)
Source repository
Open source repository

Source-reported description

The fullest retained description is shown with its source. Distribution packaging descriptions may include downstream details.

Provides an implementation of today's most used tokenizers, with a focus on performance and versatility. * Train new vocabularies and tokenize, using today's most used tokenizers. * Extremely fast (both training and tokenization), thanks to the Rust implementation. Takes less than 20 seconds to tokenize a GB of text on a server's CPU. * Easy to use, but also extremely versatile. * Designed for research and production. * Normalization comes with alignments tracking. It's always possible to get the part of the original sentence that corresponds to a given token. * Does all the pre-processing: Truncate, Pad, add the special tokens your model needs.

Description source

Packages by operating system

Compare recorded versions, then open a package for dependency, file, checksum, and repository evidence. Version strings are distribution-specific, not a ranking of newer software.

Fedora 43

  1. python3-tokenizers

    Fedora 43 / Unspecified / source python-tokenizers

    0.22.2-2.fc43

    Implementation of today's most used tokenizers

    aarch64x86_6443

Fedora 44

  1. python3-tokenizers

    Fedora 44 / Unspecified / source python-tokenizers

    0.22.2-2.fc44

    Implementation of today's most used tokenizers

    aarch64x86_6444

openSUSE Leap 16.0

  1. python313-tokenizers

    openSUSE Leap 16.0 / Unspecified / source python-tokenizers

    0.21.1-bp160.1.11

    Provides an implementation of today's most used tokenizers

    aarch64x86_64leap-16.0

openSUSE Tumbleweed

  1. python313-tokenizers

    openSUSE Tumbleweed / Unspecified / source python-tokenizers

    0.23.1-1.2

    Provides an implementation of today's most used tokenizers

    x86_64tumbleweed
  2. python314-tokenizers

    openSUSE Tumbleweed / Unspecified / source python-tokenizers

    0.23.1-1.2

    Provides an implementation of today's most used tokenizers

    x86_64tumbleweed

Project resources and further reading

Mapping provenance

Only source-backed identity signals create public cross-OS links. A reviewer can later approve or dispute an inferred relationship without rewriting native package history.

No field-level source record is published yet.