Linux workstation

openSUSE Tumbleweed native package

perl-HTML-TableExtract

Perl module for extracting the content contained in tables within an HTM[cut]

Packages / openSUSE Tumbleweed / Development/Libraries/Perl / perl-HTML-TableExtract

[Source: perl-HTML-TableExtract]

Package: perl-HTML-TableExtract (2.15-2.8)

[Project overview: libhtml-tableextract-perl]

External Resources:

Homepage: [metacpan.org]

Perl module for extracting the content contained in tables within an HTM[cut]

HTML::TableExtract is a subclass of HTML::Parser that serves to extract the information from tables of interest contained within an HTML document. The information from each extracted table is stored in table objects. Tables can be extracted as text, HTML, or HTML::ElementTable structures (for in-place editing or manipulation). There are currently four constraints available to specify which tables you would like to extract from a document: _Headers_, _Depth_, _Count_, and _Attributes_. _Headers_, the most flexible and adaptive of the techniques, involves specifying text in an array that you expect to appear above the data in the tables of interest. Once all headers have been located in a row of that table, all further cells beneath the columns that matched your headers are extracted. All other columns are ignored: think of it as vertical slices through a table. In addition, TableExtract automatically rearranges each row in the same order as the headers you provided. If you would like to disable this, set _automap_ to 0 during object creation, and instead rely on the column_map() method to find out the order in which the headers were found. Furthermore, TableExtract will automatically compensate for cell span issues so that columns are really the same columns as you would visually see in a browser. This behavior can be disabled by setting the _gridmap_ parameter to 0. HTML is stripped from the entire textual content of a cell before header matches are attempted -- unless the _keep_html_ parameter was enabled. _Depth_ and _Count_ are more specific ways to specify tables in relation to one another. _Depth_ represents how deeply a table resides in other tables. The depth of a top-level table in the document is 0. A table within a top-level table has a depth of 1, and so on. Each depth can be thought of as a layer; tables sharing the same depth are on the same layer. Within each of these layers, _Count_ represents the order in which a table was seen at that depth, starting with 0. Providing both a _depth_ and a _count_ will uniquely specify a table within a document. _Attributes_ match based on the attributes of the html <table> tag, for example, border widths or background color. Each of the _Headers_, _Depth_, _Count_, and _Attributes_ specifications are cumulative in their effect on the overall extraction. For instance, if you specify only a _Depth_, then you get all tables at that depth (note that these could very well reside in separate higher- level tables throughout the document since depth extends across tables). If you specify only a _Count_, then the tables at that _Count_ from all depths are returned (i.e., the _n_th occurrence of a table at each depth). If you only specify _Headers_, then you get all tables in the document containing those column headers. If you have specified multiple constraints of _Headers_, _Depth_, _Count_, and _Attributes_, then each constraint has veto power over whether a particular table is extracted. If no _Headers_, _Depth_, _Count_, or _Attributes_ are specified, then all tables match. When extracting only text from tables, the text is decoded with HTML::Entities by default; this can be disabled by setting the _decode_ parameter to 0.

Other Packages Related to perl-HTML-TableExtract:

  • dep: perl(:MODULE_COMPAT_5.44.0)

    Package not available

  • dep: perl(HTML::ElementTable) (>= 1.16)

    Package not available

  • dep: perl(HTML::Parser)

    Package not available

Download perl-HTML-TableExtract

ArchitecturePackage SizeInstalled SizeFiles
noarch47 KiB104 KiB[list of files]

Шляхи файлів пакета (9)

Paths come from the repository package-file index for the observed builds. They describe archive/package associations, not every file that will exist on a running system after maintainer scripts, alternatives, generated state, diversions, or installation choices.

  • /usr/lib/perl5/vendor_perl/5.44.0/HTML
  • /usr/lib/perl5/vendor_perl/5.44.0/HTML/TableExtract.pm
  • /usr/share/doc/packages/perl-HTML-TableExtract
  • /usr/share/doc/packages/perl-HTML-TableExtract/Changes
  • /usr/share/doc/packages/perl-HTML-TableExtract/examples.html
  • /usr/share/doc/packages/perl-HTML-TableExtract/README
  • /usr/share/licenses/perl-HTML-TableExtract
  • /usr/share/licenses/perl-HTML-TableExtract/LICENSE
  • /usr/share/man/man3/HTML::TableExtract.3pm.gz

Field source: openSUSE Tumbleweed OSS revision tumbleweed-oss-multiarch:4db10cdde3ad82c7e22be35753c9e164b4bdd30f92db4aee17a0d4c94c5f902f

Використати цей пакет

OpenFactory може завантажити цю операційну систему у віртуальній машині браузера або почати збірку образу з рідною назвою пакета з цього запису.

Версії, набори та репозиторії

Кожен рядок: метадані індексу пакетів для однієї версії, архітектури, набору й репозиторію. Назви, URL і розміри зі джерела; посилання є змінним місцем отримання, не перерозповсюдженням OpenFactory.

VersionReleaseArchitectureRepositoryPackage sizeInstalled sizePublisher repository artifact
2.15-2.8tumbleweed / ossnoarchopenSUSE Tumbleweed · OSS · multi-architecture47 KiB104 KiBnoarch/perl-HTML-TableExtract-2.15-2.8.noarch.rpm

Field source: openSUSE Tumbleweed OSS revision tumbleweed-oss-multiarch:4db10cdde3ad82c7e22be35753c9e164b4bdd30f92db4aee17a0d4c94c5f902f

Контрольні суми й дати спостереження

For an APT source, signature verification authenticates the repository metadata chain and the Packages index containing this source-reported artifact digest. It does not certify package safety.

Повнота запису каталогу

The completeness score measures metadata coverage, not software quality, security, compatibility, or suitability.

Summary and description
25/25
Artifact path and source digest
25/25
Dependency metadata
15/15
Package-file index
15/15
Homepage
5/5
License text
5/5
Source package or maintainer
10/10

Recorded total: 100/100

Field source: openSUSE Tumbleweed OSS revision tumbleweed-oss-multiarch:4db10cdde3ad82c7e22be35753c9e164b4bdd30f92db4aee17a0d4c94c5f902f. The cross-OS mapping is catalog-derived from the source-reported homepage; it does not establish authorship or publisher identity

Джерела та походження

Field-source links above resolve here. Each source entry names the metadata publisher, trust tier, exact snapshot revision, signature result, and observation time; catalog-derived mappings are labeled separately.

  • Authoritative source; repository metadata signature verified, revision tumbleweed-oss-multiarch:4db10cdde3ad82c7e22be35753c9e164b4bdd30f92db4aee17a0d4c94c5f902f

    Signature verification covers the configured repository metadata chain. It does not certify that the package is safe or suitable.

    Repository-signature verification record
    Signed-object SHA-256
    4db10cdde3ad82c7e22be35753c9e164b4bdd30f92db4aee17a0d4c94c5f902f
    Signer fingerprint
    AD485664E901B867051AB15F35A2F86E29B700A4
    Keyring revision
    opensuse-project-signing-key.asc
    SHA-256 b5745739ebfb95b25b8e810f9bcb847fe750ccda598bd85f02f0e974599a6d7e
    Tool and policy
    gpgv (GnuPG) 2.4.9
    openfactory-software-catalog-signature-v1
    Verification time
    Sep 1, 2026
    Signed Release → package-index hash linkage

    Path: Not recorded
    Expected SHA-256: Not recorded
    Observed SHA-256: Not recorded
    Result: match verified

perl-HTML-TableExtract Package for openSUSE Tumbleweed | OpenFactory