Recommender
The recommender is a classical meta-learner: given a dataset's meta-feature
vector, it predicts a ranking over the circuit pool. Most users only need
load_default_recommender, which returns the pre-trained bundle shipped
with the package. get_recommender builds a fresh (unfitted) recommender
for retraining; the model-selection utilities below drive the offline
search over classifiers and feature subsets.
Qmes.load_default_recommender
Load the pre-trained recommender bundled with the Qmes package.
This is the entry point for the package's core value proposition:
pip install Qmes -> recommend a circuit -> no quantum evaluation
needed at inference time. The bundle is shipped as package data
(see [tool.setuptools.package-data] in pyproject.toml) under
Qmes/_models/
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
task_type
|
(classification, regression)
|
|
'classification'
|
Returns:
| Type | Description |
|---|---|
PairwiseRecommender, already fitted, ready for .predict() or for
|
|
Qmes.recommend(X, y, extractor, recommender=...).
|
|
Raises:
| Type | Description |
|---|---|
ValueError : unknown task_type
|
|
FileNotFoundError : package was installed without the bundled model
|
data (e.g. a from-source checkout where Qmes/_models/ |
Qmes.get_recommender
get_recommender(
task_type,
classifier,
feature_indices=None,
tied_threshold=TIED_THRESHOLD,
feature_names=None,
metric_name=None,
)
Build a PairwiseRecommender with task_type/metric_name auto-filled.
There is only one recommender class (PairwiseRecommender is task-
agnostic: behavior depends on the classifier and data passed in, not
on task_type), so this factory exists mainly to remove a manual-typo
failure mode -- passing the wrong metric_name for a task_type used to
be possible by hand and would only surface as a confusing label later.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
task_type
|
(classification, regression)
|
|
'classification'
|
classifier
|
sklearn estimator (unfitted template), still supplied by
|
the caller -- there is no per-task default classifier. |
required |
feature_indices
|
list[int] or None
|
Which meta-feature columns the recommender uses. None = all. |
None
|
tied_threshold
|
float
|
Tie tolerance for downstream evaluation; stored and persisted on the recommender, not applied by predict() itself. |
TIED_THRESHOLD
|
feature_names
|
list[str] or None
|
Full ordered meta-feature names expected from the extractor (before feature_indices subsetting). Persisted so inference can assert order alignment after load(). |
None
|
metric_name
|
str or None
|
Override the auto-filled default ('MCC' for classification, 'R2' for regression) if needed. |
None
|
Returns:
| Type | Description |
|---|---|
PairwiseRecommender (unfitted; call .fit(X, pivot) before saving)
|
|
Qmes.PairwiseRecommender
PairwiseRecommender(
classifier,
feature_indices=None,
tied_threshold=TIED_THRESHOLD,
feature_names=None,
task_type=None,
metric_name=None,
)
Pairwise one-vs-one (OvO) circuit recommender.
Trains one binary comparator per circuit pair (21 comparators for the
default 7-circuit pool), each an independent clone of the base
classifier. Each comparator learns, from a dataset's
meta-features, which of its two circuits scores higher. At prediction
time every pair casts one vote and circuits are ranked by total
vote count.
A single class serves both classification and regression: behavior is
determined by the classifier template and the training data, not
by task_type (which is stored for validation against the extractor
at inference time).
Typical usage:
load_default_recommender(task)- pre-trained bundle, ready to use.get_recommender(task, clf, ...)thenfit()- retrain your own.
Attributes:
| Name | Type | Description |
|---|---|---|
circuits_ |
list[str]
|
Circuit names, set after |
pairs_ |
list[tuple[str, str]]
|
All circuit pairs, set after |
classifiers_ |
dict[tuple[str, str], object]
|
Fitted comparator (a clone of the base |
is_fitted_ |
bool
|
True once |
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
classifier
|
sklearn estimator (unfitted template)
|
|
required |
feature_indices
|
list[int] or None
|
Which meta-feature columns to use. None = all. |
None
|
tied_threshold
|
float
|
Tie tolerance for downstream evaluation (run_loo_evaluation, evaluate_recommendation). |
TIED_THRESHOLD
|
feature_names
|
list[str] or None
|
Full ordered list of meta-feature names this recommender expects from the extractor (before feature_indices subsetting). Persisted so inference can assert order alignment after load(). |
None
|
task_type
|
str or None
|
E.g. 'classification' or 'regression'. Persisted so inference can assert the recommender matches the extractor/evaluator in use. |
None
|
metric_name
|
str or None
|
E.g. 'MCC' or 'R2'. Stored for traceability only. |
None
|
fit
Train one comparator per circuit pair.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
meta_features
|
(n_datasets, d) meta-feature matrix
|
Row i must correspond to pivot_scores.columns[i]. Alignment is positional - if a DataFrame is passed, only its .values are used, not its index. |
required |
pivot_scores
|
DataFrame, index=circuits, columns=datasets
|
Values = primary metric (MCC, R², etc.) |
required |
predict
Predict circuit ranking for a new dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
meta_features
|
(d,) or (1, d) meta-feature vector
|
|
required |
top_k
|
number of top circuits to return
|
|
3
|
Returns:
| Type | Description |
|---|---|
dict with keys:
|
'ranking': list of all circuits sorted by votes (descending) 'top_k': list of top-k circuits 'votes': dict {circuit: vote_count} |
save
Save the recommender to a directory as a format-v2 bundle.
Writes two files: recommender.npz (training meta-feature
matrix and pivot score values) and recommender.json
(classifier spec and configuration). Fitted estimators are NOT
persisted; load() refits from the stored training data instead,
so the bundle carries no pickled sklearn objects and stays
independent of the sklearn version that produced it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str or Path
|
Target directory. Created if it does not exist. |
required |
Raises:
| Type | Description |
|---|---|
RuntimeError
|
If called before fit(). |
TypeError
|
If the classifier's get_params() contains values that are not JSON-serializable (e.g. estimator objects). |
load
classmethod
Load a format-v2 bundle and refit from its stored training data.
The classifier is reconstructed from the spec in
recommender.json and refit on the training matrices in
recommender.npz. Refitting is cheap and deterministic for
the bundled kNN (no random state), and removes any coupling to
the sklearn version that created the bundle.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str or Path
|
Directory containing |
required |
Returns:
| Type | Description |
|---|---|
PairwiseRecommender
|
A fitted recommender equivalent to the one that was saved. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the directory contains a legacy format-v1 pickle bundle. |
FileNotFoundError
|
If the bundle files are missing. |
Model selection utilities
Qmes.run_loo_evaluation
run_loo_evaluation(
meta_features,
pivot_scores,
classifiers=None,
feature_subsets=None,
tied_threshold=TIED_THRESHOLD,
verbose=True,
dataset_names=None,
)
Exhaustive LOO evaluation over (classifier × feature_subset).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
meta_features
|
(n_datasets, d)
|
Row i MUST correspond to |
required |
pivot_scores
|
index=circuits, columns=datasets
|
|
required |
classifiers
|
dict {name: sklearn_estimator}
|
|
None
|
feature_subsets
|
dict {label: list_of_indices}
|
|
None
|
tied_threshold
|
for tied-best evaluation
|
|
TIED_THRESHOLD
|
verbose
|
print progress
|
|
True
|
dataset_names
|
list[str] or None
|
Optional alignment guard. If given, must be the dataset name for each
meta-feature row IN ROW ORDER; the function asserts it equals
|
None
|
Returns:
| Type | Description |
|---|---|
DataFrame with columns: Features, Classifier, Single, Tied, Top3_Tied, Mean_Regret
|
|
Qmes.recommender.select_features_mi
Rank meta-features by mutual information and return index subsets.
MI is computed between each meta-feature and labels (the best circuit
per dataset - a categorical target), then features are ranked
descending. Used at training time to test how few meta-features still
yield a good recommender.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
meta_features
|
(ndarray, shape(n_datasets, d))
|
Meta-feature matrix, one row per dataset. |
required |
labels
|
(ndarray, shape(n_datasets))
|
Categorical target - the best circuit for each dataset, e.g.
|
required |
k_values
|
tuple[int, ...]
|
Subset sizes to produce. A size is skipped if it exceeds d. |
(5, 10, 15, 20)
|
random_state
|
int
|
Seed for mutual_info_classif (its k-NN estimator is stochastic). |
42
|
Returns:
| Type | Description |
|---|---|
dict[str, list[int]]
|
Index subsets keyed by label: {"full": all d indices, "top5": [...], "top10": [...], ...}, each a list of column indices into meta_features ordered by descending MI. |
Qmes.recommender.DEFAULT_CLASSIFIERS
module-attribute
DEFAULT_CLASSIFIERS = {
"DT": DecisionTreeClassifier(random_state=42),
"RF": RandomForestClassifier(
n_estimators=100, random_state=42
),
"GB": GradientBoostingClassifier(random_state=42),
"AdaBoost": AdaBoostClassifier(
estimator=DecisionTreeClassifier(random_state=42),
random_state=42,
),
"Bagging": BaggingClassifier(random_state=42),
"SVM-linear": SVC(
kernel="linear", probability=True, random_state=42
),
"SVM-rbf": SVC(
kernel="rbf", probability=True, random_state=42
),
"SVM-sigmoid": SVC(
kernel="sigmoid", probability=True, random_state=42
),
"MLP-small": MLPClassifier(
hidden_layer_sizes=(64, 32),
max_iter=500,
random_state=42,
),
"MLP-large": MLPClassifier(
hidden_layer_sizes=(128, 64, 32),
max_iter=500,
random_state=42,
),
"kNN": KNeighborsClassifier(),
"NearCentroid": NearestCentroid(),
"NaiveBayes": GaussianNB(),
"LogReg": LogisticRegression(
max_iter=1000, random_state=42
),
}
The 14-classifier pool searched by run_loo_evaluation.
Spans tree-based, ensemble, SVM, neural-network, instance-based, and
probabilistic families so the model-selection grid covers diverse
inductive biases. All stochastic members are seeded (random_state=42)
for reproducible LOO results. The shipped default recommenders use kNN,
selected from this pool by LOO mean regret.