Machine learning
scikit-learn datasets, models, evaluation and tuning.
15 nodes. Right-click any node in the editor to read this documentation in the app.
ML / Dataโ
๐ Apply Transformerโ
id sklearn_transform ยท ML / Data ยท Python export: yes
Fit a transformer on the TRAINING rows only, then transform both training and test rows (so nothing about the test data leaks into preprocessing). Outputs a new splits dict; y is unchanged.
Inputs
| Port | Type | Description |
|---|---|---|
transformer | Transformer | Unfitted transformer from a Transformer node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
splits | Splits | Splits with X_train / X_test transformed. |
๐ธ Sample Datasetโ
id sklearn_load_dataset ยท ML / Data ยท Python export: yes
One of scikit-learn's built-in datasets as a DataFrame: the feature columns plus a 'target' column. iris, wine, breast_cancer and digits are classification problems; diabetes is regression. Ships with scikit-learn, so it works offline.
Outputs
| Port | Type | Description |
|---|---|---|
dataframe | DataFrame | Feature columns plus a 'target' column. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Dataset | select | iris | iris, wine, breast_cancer, diabetes, digits |
โ๏ธ Train/Test Splitโ
id sklearn_train_test_split ยท ML / Data ยท Python export: yes
Split a DataFrame into training and test sets. Name the target column; every other column becomes a feature. A fixed random seed makes the split reproducible (use a negative number for a different split each run); 'Stratify' keeps class proportions equal in both sets.
Inputs
| Port | Type | Description |
|---|---|---|
dataframe | DataFrame | Table with the features and the target column. |
Outputs
| Port | Type | Description |
|---|---|---|
splits | Splits | Dict of X_train, X_test, y_train, y_test (pandas objects). |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Target column | text | target | |
| Test fraction | float | 0.2 | |
| Random seed (negative = random) | int | 42 | |
| Stratify by target | checkbox | false | |
| Shuffle before splitting | checkbox | true |
๐ง Transformerโ
id sklearn_transformer ยท ML / Data ยท Python export: yes
Configure a preprocessing step: standard / min-max / robust scaling, missing-value imputation (mean, median, most frequent), or PCA. Nothing is fitted here โ wire it into Apply Transformer. 'Components' applies to PCA only; 'Extra parameters' is an optional JSON object of further scikit-learn arguments.
Outputs
| Port | Type | Description |
|---|---|---|
transformer | Transformer | Unfitted scikit-learn transformer. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Transformer | select | standard_scaler | standard_scaler, minmax_scaler, robust_scaler, impute_mean, impute_median, impute_most_frequent, pca |
| Components (PCA only) | int | 2 | |
| Extra parameters (JSON) | text |
ML / Evaluationโ
๐ฒ Confusion Matrixโ
id sklearn_confusion_matrix ยท ML / Evaluation ยท Python export: yes
For a fitted classifier: how often each actual class was predicted as each class (rows = actual, columns = predicted). The diagonal is the correct predictions.
Inputs
| Port | Type | Description |
|---|---|---|
model | FittedModel | A fitted model โ from Fit or a Hyper-parameter Search node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
matrix | DataFrame | Counts, rows = actual class, columns = predicted class. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Rows | select | test | test, train |
๐ Cross-Validateโ
id sklearn_cross_val ยท ML / Evaluation ยท Python export: yes
k-fold cross-validation of an (unfitted) estimator on the TRAINING rows: a more reliable estimate than a single split. Outputs one score per fold; scoring 'default' uses the model's own (accuracy for classifiers, Rยฒ for regressors).
Inputs
| Port | Type | Description |
|---|---|---|
estimator | Estimator | Unfitted model from a Classifier, Regressor or Clusterer node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
scores | DataFrame | One row per fold: fold, score. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Folds | int | 5 | |
| Scoring | select | default | default, accuracy, f1_weighted, roc_auc, r2, neg_mean_squared_error, neg_mean_absolute_error |
๐ Evaluateโ
id sklearn_evaluate ยท ML / Evaluation ยท Python export: yes
Score a fitted model. Classifiers: accuracy, precision, recall and F1 (class-weighted) plus ROC-AUC when the model gives probabilities. Regressors: Rยฒ, MAE, MSE and RMSE. Outputs a metric / value table (plot it with Bar Chart).
Inputs
| Port | Type | Description |
|---|---|---|
model | FittedModel | A fitted model โ from Fit or a Hyper-parameter Search node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
metrics | DataFrame | Two columns: metric, value. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Rows | select | test | test, train |
โญ Feature Importanceโ
id sklearn_feature_importance ยท ML / Evaluation ยท Python export: yes
Which features drive a fitted model? Uses the model's feature_importances_ (forests, boosting, trees) or the magnitude of its coefficients (linear models, averaged over classes). Sorted, most important first.
Inputs
| Port | Type | Description |
|---|---|---|
model | FittedModel | A fitted model โ from Fit or a Hyper-parameter Search node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
importance | DataFrame | Columns: feature, importance. |
ML / Modelsโ
๐ง Classifierโ
id sklearn_classifier ยท ML / Models ยท Python export: yes
Configure a classification model: logistic regression, random forest, gradient boosting, SVM, k-NN or decision tree. Nothing is trained until Fit. Only the parameters that apply to the chosen model are used: 'Trees' โ random forest / gradient boosting; 'Max depth' โ forests, boosting, decision tree (0 = no limit); 'C' โ logistic regression / SVM; 'Neighbours' โ k-NN; 'Learning rate' โ gradient boosting; 'Kernel' โ SVM. 'Extra parameters' is an optional JSON object of further scikit-learn arguments, e.g. {"min_samples_leaf": 3}.
Outputs
| Port | Type | Description |
|---|---|---|
estimator | Estimator | Unfitted scikit-learn model โ wire into Fit, Cross-Validate or Search. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Model | select | random_forest | logistic_regression, random_forest, gradient_boosting, svm, knn, decision_tree |
| Trees | int | 100 | |
| Max depth (0 = no limit) | int | 0 | |
| C (regularisation) | float | 1.0 | |
| Neighbours | int | 5 | |
| Learning rate | float | 0.1 | |
| Kernel | select | rbf | rbf, linear, poly, sigmoid |
| Random seed (negative = random) | int | 42 | |
| Extra parameters (JSON) | text |
๐ซง Clustererโ
id sklearn_clusterer ยท ML / Models ยท Python export: yes
Configure an unsupervised clustering model: k-means, DBSCAN, agglomerative or Gaussian mixture. 'Clusters' applies to k-means, agglomerative and the mixture; 'Eps' / 'Min samples' to DBSCAN. Fit it on the features (labels are ignored), then use Predict.
Outputs
| Port | Type | Description |
|---|---|---|
estimator | Estimator | Unfitted scikit-learn model โ wire into Fit, Cross-Validate or Search. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Algorithm | select | kmeans | kmeans, dbscan, agglomerative, gaussian_mixture |
| Clusters | int | 3 | |
| Eps (DBSCAN) | float | 0.5 | |
| Min samples (DBSCAN) | int | 5 | |
| Random seed (negative = random) | int | 42 | |
| Extra parameters (JSON) | text |
๐๏ธ Fitโ
id sklearn_fit ยท ML / Models ยท Python export: yes
Train a model on the training rows (X_train, and y_train when the model is supervised). Works on a copy, so the estimator node can feed other nodes unchanged. Outputs the fitted model.
Inputs
| Port | Type | Description |
|---|---|---|
estimator | Estimator | Unfitted model from a Classifier, Regressor or Clusterer node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
model | FittedModel | The trained model โ wire into Predict, Evaluate, Feature Importance. |
๐ฎ Predictโ
id sklearn_predict ยท ML / Models ยท Python export: yes
Run a fitted model on the test (or training) rows and output its predictions as a Series indexed like the rows. For clusterers these are the cluster labels; DBSCAN and agglomerative clustering can only label their training rows.
Inputs
| Port | Type | Description |
|---|---|---|
model | FittedModel | A fitted model โ from Fit or a Hyper-parameter Search node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
predictions | Series | One prediction per row, named 'prediction'. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Rows | select | test | test, train |
๐ Regressorโ
id sklearn_regressor ยท ML / Models ยท Python export: yes
Configure a regression model: linear, ridge, lasso, random forest, gradient boosting, SVR, k-NN or decision tree. Nothing is trained until Fit. 'Alpha' applies to ridge / lasso. Only the parameters that apply to the chosen model are used: 'Trees' โ random forest / gradient boosting; 'Max depth' โ forests, boosting, decision tree (0 = no limit); 'C' โ logistic regression / SVM; 'Neighbours' โ k-NN; 'Learning rate' โ gradient boosting; 'Kernel' โ SVM. 'Extra parameters' is an optional JSON object of further scikit-learn arguments, e.g. {"min_samples_leaf": 3}.
Outputs
| Port | Type | Description |
|---|---|---|
estimator | Estimator | Unfitted scikit-learn model โ wire into Fit, Cross-Validate or Search. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Model | select | random_forest | linear_regression, ridge, lasso, random_forest, gradient_boosting, svr, knn, decision_tree |
| Alpha (ridge / lasso) | float | 1.0 | |
| Trees | int | 100 | |
| Max depth (0 = no limit) | int | 0 | |
| Neighbours | int | 5 | |
| Learning rate | float | 0.1 | |
| C (SVR) | float | 1.0 | |
| Kernel (SVR) | select | rbf | rbf, linear, poly, sigmoid |
| Random seed (negative = random) | int | 42 | |
| Extra parameters (JSON) | text |
ML / Tuningโ
๐๏ธ Hyper-parameter Searchโ
id sklearn_search ยท ML / Tuning ยท Python export: yes
Find the best settings by cross-validation on the TRAINING rows. Give a JSON grid of parameter names to lists of values, e.g. {"n_estimators": [50, 100], "max_depth": [3, 5, 10]} โ use the scikit-learn parameter names of the estimator. 'grid' tries every combination; 'random' tries a random sample of them. The output is the fitted search: it acts as the best model (wire it into Predict / Evaluate) and into Search Results for the full table.
Inputs
| Port | Type | Description |
|---|---|---|
estimator | Estimator | Unfitted model from a Classifier, Regressor or Clusterer node. |
splits | Splits | Train/test data โ from Train/Test Split or Apply Transformer. |
Outputs
| Port | Type | Description |
|---|---|---|
search | FittedModel | Fitted search; predicts with the best parameters found. |
Fields
| Field | Type | Default | Choices |
|---|---|---|---|
| Strategy | select | grid | grid, random |
| Parameter grid (JSON) | textarea | {"n_estimators": [50, 100], "max_depth": [3, 5,โฆ | |
| CV folds | int | 5 | |
| Scoring | select | default | default, accuracy, f1_weighted, roc_auc, r2, neg_mean_squared_error, neg_mean_absolute_error |
| Samples (random strategy) | int | 10 | |
| Random seed (negative = random) | int | 42 |
๐ Search Resultsโ
id sklearn_search_results ยท ML / Tuning ยท Python export: yes
Every parameter combination a search tried, with its mean and spread of cross-validation score and rank โ best first.
Inputs
| Port | Type | Description |
|---|---|---|
search | FittedModel | Output of a Hyper-parameter Search node. |
Outputs
| Port | Type | Description |
|---|---|---|
results | DataFrame | One row per combination: parameters, mean_test_score, std_test_score, rank_test_score. |