Linux workstation

Upstream software project

libhtml-tableextract-perl

Perl module for extracting the content contained in tables within an HTM[cut]

About libhtml-tableextract-perl

Perl module for extracting the content contained in tables within an HTM[cut]

This project links 4 native package records across 4 recorded operating-system releases. Compare the retained versions and architectures below, then open the package for your own release.

These are catalog observations, not a guarantee of installation, compatibility, or upstream support.

Project pictures and package coverage

Debian 12 (Bookworm): 1 package records; Debian 13 (Trixie): 1 package records; openSUSE Leap 15.6: 1 package records; openSUSE Tumbleweed: 1 package records. Catalog coverage diagram, not an application screenshot.libhtml-tableextract-perl: recorded package coverageDebian 12 (Bookworm)1 recordsDebian 13 (Trixie)1 recordsopenSUSE Leap 15.61 recordsopenSUSE Tumbleweed1 records
OpenFactory diagram of linked package records. It is not an application screenshot.

Project identity

Project
libhtml-tableextract-perl
Publisher
Not authoritatively mapped
Native package records
4
Operating systems
debian-12, debian-13, opensuse-leap-15-6, opensuse-tumbleweed
License expression
Artistic-1.0 OR GPL-1.0-or-later
Metadata completeness
100/100 (not a software quality rating)
Source repository
Not reported

Source-reported description

The fullest retained description is shown with its source. Distribution packaging descriptions may include downstream details.

HTML::TableExtract is a subclass of HTML::Parser that serves to extract the information from tables of interest contained within an HTML document. The information from each extracted table is stored in table objects. Tables can be extracted as text, HTML, or HTML::ElementTable structures (for in-place editing or manipulation). There are currently four constraints available to specify which tables you would like to extract from a document: _Headers_, _Depth_, _Count_, and _Attributes_. _Headers_, the most flexible and adaptive of the techniques, involves specifying text in an array that you expect to appear above the data in the tables of interest. Once all headers have been located in a row of that table, all further cells beneath the columns that matched your headers are extracted. All other columns are ignored: think of it as vertical slices through a table. In addition, TableExtract automatically rearranges each row in the same order as the headers you provided. If you would like to disable this, set _automap_ to 0 during object creation, and instead rely on the column_map() method to find out the order in which the headers were found. Furthermore, TableExtract will automatically compensate for cell span issues so that columns are really the same columns as you would visually see in a browser. This behavior can be disabled by setting the _gridmap_ parameter to 0. HTML is stripped from the entire textual content of a cell before header matches are attempted -- unless the _keep_html_ parameter was enabled. _Depth_ and _Count_ are more specific ways to specify tables in relation to one another. _Depth_ represents how deeply a table resides in other tables. The depth of a top-level table in the document is 0. A table within a top-level table has a depth of 1, and so on. Each depth can be thought of as a layer; tables sharing the same depth are on the same layer. Within each of these layers, _Count_ represents the order in which a table was seen at that depth, starting with 0. Providing both a _depth_ and a _count_ will uniquely specify a table within a document. _Attributes_ match based on the attributes of the html <table> tag, for example, border widths or background color. Each of the _Headers_, _Depth_, _Count_, and _Attributes_ specifications are cumulative in their effect on the overall extraction. For instance, if you specify only a _Depth_, then you get all tables at that depth (note that these could very well reside in separate higher- level tables throughout the document since depth extends across tables). If you specify only a _Count_, then the tables at that _Count_ from all depths are returned (i.e., the _n_th occurrence of a table at each depth). If you only specify _Headers_, then you get all tables in the document containing those column headers. If you have specified multiple constraints of _Headers_, _Depth_, _Count_, and _Attributes_, then each constraint has veto power over whether a particular table is extracted. If no _Headers_, _Depth_, _Count_, or _Attributes_ are specified, then all tables match. When extracting only text from tables, the text is decoded with HTML::Entities by default; this can be disabled by setting the _decode_ parameter to 0.

Description source

Packages by operating system

Compare recorded versions, then open a package for dependency, file, checksum, and repository evidence. Version strings are distribution-specific, not a ranking of newer software.

Debian 12 (Bookworm)

  1. libhtml-tableextract-perl

    Debian 12 (Bookworm) / perl

    2.15-2

    module for extracting the content contained in HTML tables

    allbookworm

Debian 13 (Trixie)

  1. libhtml-tableextract-perl

    Debian 13 (Trixie) / perl

    2.15-2

    module for extracting the content contained in HTML tables

    alltrixie

openSUSE Leap 15.6

  1. perl-HTML-TableExtract

    openSUSE Leap 15.6 / Development/Libraries/Perl / source perl-HTML-TableExtract

    2.15-bp156.3.1

    Perl module for extracting the content contained in tables within an HTM[cut]

    noarchleap-15.6

openSUSE Tumbleweed

  1. perl-HTML-TableExtract

    openSUSE Tumbleweed / Development/Libraries/Perl / source perl-HTML-TableExtract

    2.15-2.8

    Perl module for extracting the content contained in tables within an HTM[cut]

    noarchtumbleweed

Project resources and further reading

Mapping provenance

Only source-backed identity signals create public cross-OS links. A reviewer can later approve or dispute an inferred relationship without rewriting native package history.

No field-level source record is published yet.

libhtml-tableextract-perl Software and Packages | OpenFactory