Skip to content

Extractors

An extractor turns a dataset (X, y) into a fixed-length meta-feature vector. Two concrete extractors ship with Qmes, both backed by problexity complexity measures - see Meta-features for the full list.

To write your own extractor, implement the three-member contract of BaseExtractor below (task_type, _feature_names, _extract_raw); a worked example is in Advanced Usage.

Qmes.get_extractor

get_extractor(task_type, **kwargs)

Return the extractor for the given task type.

Parameters:

Name Type Description Default
task_type (classification, regression)
'classification'
**kwargs passed to extractor constructor
{}

Returns:

Type Description
Concrete BaseExtractor instance

Qmes.extractors.BaseExtractor

Bases: ABC

Abstract base class for task-specific meta-feature extractors.

Subclasses must implement: - task_type (property): str identifier - _feature_names (property): list of feature names (fixed-length) - _extract_raw(X, y): compute meta-features, return ndarray

Lifecycle: 1. User calls extract(X, y) or extract(X) 2. Base class validates input 3. Calls _extract_raw(X, y) → raw vector 4. Sanitizes (NaN/Inf → 0.0) 5. Wraps in ExtractionResult

task_type abstractmethod property

task_type

Task type identifier, e.g. 'classification'.

_feature_names abstractmethod property

_feature_names

Fixed list of feature names. Length defines output dimension.

_extract_raw abstractmethod

_extract_raw(X, y=None)

Compute raw meta-feature vector.

Parameters:

Name Type Description Default
X feature matrix or time series matrix
required
y target vector
None

Returns:

Type Description
ndarray shape (d,) where d == len(self._feature_names)

extract

extract(X, y=None)

Public API: extract meta-features from dataset.

Parameters:

Name Type Description Default
X (ndarray, shape(n_samples, n_features))
required
y (ndarray, shape(n_samples))

Target vector. Required for both classification and regression.

None

Returns:

Type Description
an ExtractionResult (vector + feature_names + task_type)

extract_batch

extract_batch(datasets)

Extract meta-features for all datasets, returns a DataFrame.

Parameters:

Name Type Description Default
datasets dict[name, data]

Output from data loader. Values are tuples (X, y) for supervised tasks. A bare ndarray (no target) is accepted by this method's signature as a forward-compatibility hook for a future unsupervised extractor, but neither extractor currently shipped (ClassificationExtractor, RegressionExtractor) supports it - both raise ValueError if y is None, so a bare-ndarray entry is caught, logged, and skipped rather than silently succeeding.

required

Returns:

Type Description
DataFrame shape (n_datasets, d)

Index = dataset names, columns = feature names. Follows the meta-dataset format from the paper:

f1 f1v ... n4
Blobs_F2C2_S100 0.0068 0.0020 ... 0.9861
Iris_S80 0.0254 0.0079 ... 0.3643

Qmes.extractors.ExtractionResult dataclass

ExtractionResult(vector, feature_names, task_type)

Container for extraction results.

Attributes:

Name Type Description
vector ndarray shape (d,), float64

Meta-feature vector. Guaranteed: no NaN, no Inf.

feature_names list[str]

Feature names, len == len(vector).

task_type str

Task type identifier (e.g. "classification", "regression").

to_dict

to_dict()

Return {feature_name: value} mapping.

Qmes.ClassificationExtractor

Bases: BaseExtractor

Meta-feature extractor for tabular classification.

Produces 22-dim vector

Qmes.RegressionExtractor

Bases: BaseExtractor

Meta-feature extractor for tabular regression.

Produces 12-dim vector