Skip to content

Recommender

The recommender is a classical meta-learner: given a dataset's meta-feature vector, it predicts a ranking over the circuit pool. Most users only need load_default_recommender, which returns the pre-trained bundle shipped with the package. get_recommender builds a fresh (unfitted) recommender for retraining; the model-selection utilities below drive the offline search over classifiers and feature subsets.

Qmes.load_default_recommender

load_default_recommender(task_type)

Load the pre-trained recommender bundled with the Qmes package.

This is the entry point for the package's core value proposition: pip install Qmes -> recommend a circuit -> no quantum evaluation needed at inference time. The bundle is shipped as package data (see [tool.setuptools.package-data] in pyproject.toml) under Qmes/_models//, not regenerated at import time.

Parameters:

Name Type Description Default
task_type (classification, regression)
'classification'

Returns:

Type Description
PairwiseRecommender, already fitted, ready for .predict() or for
Qmes.recommend(X, y, extractor, recommender=...).

Raises:

Type Description
ValueError : unknown task_type
FileNotFoundError : package was installed without the bundled model

data (e.g. a from-source checkout where Qmes/_models// was never populated -- this is a packaging/build issue, not a normal runtime condition).

Qmes.get_recommender

get_recommender(
    task_type,
    classifier,
    feature_indices=None,
    tied_threshold=TIED_THRESHOLD,
    feature_names=None,
    metric_name=None,
)

Build a PairwiseRecommender with task_type/metric_name auto-filled.

There is only one recommender class (PairwiseRecommender is task- agnostic: behavior depends on the classifier and data passed in, not on task_type), so this factory exists mainly to remove a manual-typo failure mode -- passing the wrong metric_name for a task_type used to be possible by hand and would only surface as a confusing label later.

Parameters:

Name Type Description Default
task_type (classification, regression)
'classification'
classifier sklearn estimator (unfitted template), still supplied by

the caller -- there is no per-task default classifier.

required
feature_indices list[int] or None

Which meta-feature columns the recommender uses. None = all.

None
tied_threshold float

Tie tolerance for downstream evaluation; stored and persisted on the recommender, not applied by predict() itself.

TIED_THRESHOLD
feature_names list[str] or None

Full ordered meta-feature names expected from the extractor (before feature_indices subsetting). Persisted so inference can assert order alignment after load().

None
metric_name str or None

Override the auto-filled default ('MCC' for classification, 'R2' for regression) if needed.

None

Returns:

Type Description
PairwiseRecommender (unfitted; call .fit(X, pivot) before saving)

Qmes.PairwiseRecommender

PairwiseRecommender(
    classifier,
    feature_indices=None,
    tied_threshold=TIED_THRESHOLD,
    feature_names=None,
    task_type=None,
    metric_name=None,
)

Pairwise one-vs-one (OvO) circuit recommender.

Trains one binary comparator per circuit pair (21 comparators for the default 7-circuit pool), each an independent clone of the base classifier. Each comparator learns, from a dataset's meta-features, which of its two circuits scores higher. At prediction time every pair casts one vote and circuits are ranked by total vote count.

A single class serves both classification and regression: behavior is determined by the classifier template and the training data, not by task_type (which is stored for validation against the extractor at inference time).

Typical usage:

  • load_default_recommender(task) - pre-trained bundle, ready to use.
  • get_recommender(task, clf, ...) then fit() - retrain your own.

Attributes:

Name Type Description
circuits_ list[str]

Circuit names, set after fit() (pivot index order).

pairs_ list[tuple[str, str]]

All circuit pairs, set after fit().

classifiers_ dict[tuple[str, str], object]

Fitted comparator (a clone of the base classifier) per pair, set after fit().

is_fitted_ bool

True once fit() has run.

Parameters:

Name Type Description Default
classifier sklearn estimator (unfitted template)
required
feature_indices list[int] or None

Which meta-feature columns to use. None = all.

None
tied_threshold float

Tie tolerance for downstream evaluation (run_loo_evaluation, evaluate_recommendation).

TIED_THRESHOLD
feature_names list[str] or None

Full ordered list of meta-feature names this recommender expects from the extractor (before feature_indices subsetting). Persisted so inference can assert order alignment after load().

None
task_type str or None

E.g. 'classification' or 'regression'. Persisted so inference can assert the recommender matches the extractor/evaluator in use.

None
metric_name str or None

E.g. 'MCC' or 'R2'. Stored for traceability only.

None

fit

fit(meta_features, pivot_scores)

Train one comparator per circuit pair.

Parameters:

Name Type Description Default
meta_features (n_datasets, d) meta-feature matrix

Row i must correspond to pivot_scores.columns[i]. Alignment is positional - if a DataFrame is passed, only its .values are used, not its index.

required
pivot_scores DataFrame, index=circuits, columns=datasets

Values = primary metric (MCC, R², etc.)

required

predict

predict(meta_features, top_k=3)

Predict circuit ranking for a new dataset.

Parameters:

Name Type Description Default
meta_features (d,) or (1, d) meta-feature vector
required
top_k number of top circuits to return
3

Returns:

Type Description
dict with keys:

'ranking': list of all circuits sorted by votes (descending) 'top_k': list of top-k circuits 'votes': dict {circuit: vote_count}

save

save(path)

Save the recommender to a directory as a format-v2 bundle.

Writes two files: recommender.npz (training meta-feature matrix and pivot score values) and recommender.json (classifier spec and configuration). Fitted estimators are NOT persisted; load() refits from the stored training data instead, so the bundle carries no pickled sklearn objects and stays independent of the sklearn version that produced it.

Parameters:

Name Type Description Default
path str or Path

Target directory. Created if it does not exist.

required

Raises:

Type Description
RuntimeError

If called before fit().

TypeError

If the classifier's get_params() contains values that are not JSON-serializable (e.g. estimator objects).

load classmethod

load(path)

Load a format-v2 bundle and refit from its stored training data.

The classifier is reconstructed from the spec in recommender.json and refit on the training matrices in recommender.npz. Refitting is cheap and deterministic for the bundled kNN (no random state), and removes any coupling to the sklearn version that created the bundle.

Parameters:

Name Type Description Default
path str or Path

Directory containing recommender.npz and recommender.json.

required

Returns:

Type Description
PairwiseRecommender

A fitted recommender equivalent to the one that was saved.

Raises:

Type Description
ValueError

If the directory contains a legacy format-v1 pickle bundle.

FileNotFoundError

If the bundle files are missing.

Model selection utilities

Qmes.run_loo_evaluation

run_loo_evaluation(
    meta_features,
    pivot_scores,
    classifiers=None,
    feature_subsets=None,
    tied_threshold=TIED_THRESHOLD,
    verbose=True,
    dataset_names=None,
)

Exhaustive LOO evaluation over (classifier × feature_subset).

Parameters:

Name Type Description Default
meta_features (n_datasets, d)

Row i MUST correspond to pivot_scores.columns[i]. Alignment is positional: this function never joins by name (meta_features is a bare array with no labels). Callers are responsible for passing meta rows in pivot-column order, e.g. meta.loc[pivot.columns].values.

required
pivot_scores index=circuits, columns=datasets
required
classifiers dict {name: sklearn_estimator}
None
feature_subsets dict {label: list_of_indices}
None
tied_threshold for tied-best evaluation
TIED_THRESHOLD
verbose print progress
True
dataset_names list[str] or None

Optional alignment guard. If given, must be the dataset name for each meta-feature row IN ROW ORDER; the function asserts it equals list(pivot_scores.columns) and raises ValueError otherwise. Pass list(meta.index) to turn the positional-alignment invariant from a convention the caller must remember into a contract the library checks.

None

Returns:

Type Description
DataFrame with columns: Features, Classifier, Single, Tied, Top3_Tied, Mean_Regret

Qmes.recommender.select_features_mi

select_features_mi(
    meta_features,
    labels,
    k_values=(5, 10, 15, 20),
    random_state=42,
)

Rank meta-features by mutual information and return index subsets.

MI is computed between each meta-feature and labels (the best circuit per dataset - a categorical target), then features are ranked descending. Used at training time to test how few meta-features still yield a good recommender.

Parameters:

Name Type Description Default
meta_features (ndarray, shape(n_datasets, d))

Meta-feature matrix, one row per dataset.

required
labels (ndarray, shape(n_datasets))

Categorical target - the best circuit for each dataset, e.g. pivot_scores.idxmax(axis=0). MI is computed via mutual_info_classif, so this must be the circuit label, NOT the dataset's own y.

required
k_values tuple[int, ...]

Subset sizes to produce. A size is skipped if it exceeds d.

(5, 10, 15, 20)
random_state int

Seed for mutual_info_classif (its k-NN estimator is stochastic).

42

Returns:

Type Description
dict[str, list[int]]

Index subsets keyed by label: {"full": all d indices, "top5": [...], "top10": [...], ...}, each a list of column indices into meta_features ordered by descending MI.

Qmes.recommender.DEFAULT_CLASSIFIERS module-attribute

DEFAULT_CLASSIFIERS = {
    "DT": DecisionTreeClassifier(random_state=42),
    "RF": RandomForestClassifier(
        n_estimators=100, random_state=42
    ),
    "GB": GradientBoostingClassifier(random_state=42),
    "AdaBoost": AdaBoostClassifier(
        estimator=DecisionTreeClassifier(random_state=42),
        random_state=42,
    ),
    "Bagging": BaggingClassifier(random_state=42),
    "SVM-linear": SVC(
        kernel="linear", probability=True, random_state=42
    ),
    "SVM-rbf": SVC(
        kernel="rbf", probability=True, random_state=42
    ),
    "SVM-sigmoid": SVC(
        kernel="sigmoid", probability=True, random_state=42
    ),
    "MLP-small": MLPClassifier(
        hidden_layer_sizes=(64, 32),
        max_iter=500,
        random_state=42,
    ),
    "MLP-large": MLPClassifier(
        hidden_layer_sizes=(128, 64, 32),
        max_iter=500,
        random_state=42,
    ),
    "kNN": KNeighborsClassifier(),
    "NearCentroid": NearestCentroid(),
    "NaiveBayes": GaussianNB(),
    "LogReg": LogisticRegression(
        max_iter=1000, random_state=42
    ),
}

The 14-classifier pool searched by run_loo_evaluation.

Spans tree-based, ensemble, SVM, neural-network, instance-based, and probabilistic families so the model-selection grid covers diverse inductive biases. All stochastic members are seeded (random_state=42) for reproducible LOO results. The shipped default recommenders use kNN, selected from this pool by LOO mean regret.