A scikit-learn pipeline lets you treat preprocessing and classification as one estimator: fit the complete workflow on training data, then use that same workflow to predict on held-out data or new passengers. For Titanic data, a ColumnTransformer can impute and scale numeric fields while imputing and one-hot encoding categorical fields; a Pipeline then connects those transformations to a classifier.
Load the Titanic data and inspect its columns
The official scikit-learn example fetches the Titanic dataset from OpenML and returns feature data as a pandas DataFrame alongside the target:
As an Amazon Associate I earn from qualifying purchases.
from sklearn.datasets import fetch_openml
X, y = fetch_openml(
"titanic",
version=1,
as_frame=True,
return_X_y=True,
)
print(X.dtypes)
print(X.isna().sum())
print(y.head())
Here, X contains input features and y is the survival target, named survived in the dataset. Inspect actual column names and missing values before choosing fields: the official example uses age and fare as numeric features, and embarked, sex, and pclass as categorical features. See the official mixed-type Titanic example.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →For a compact demonstration, use those five fields. This is a feature-selection choice for the example, not a claim that they are the only useful predictors. If you change the fields, make sure the names match the fetched DataFrame.
#1 Best Overall
- TITANIC SHIP SCENES, PASSENGERS AND VINTAGE DETAILS: Color historical ocean liner illustrations featuring promenade decks, elegant travelers, mothers and children, photographers, portholes, deck chairs, luggage, ship equipment, cabins, nautical details, and Edwardian maritime scenes created for Titanic fans, history lovers, collectors, seniors, beginners, and adult colorists.
- Thick Cardstock Paper: Each design is printed on substantial cardstock for a sturdier coloring surface. The single-sided format gives every illustration its own page and helps protect the next design while coloring.
- Detailed Designs for Adults: This spiral adult coloring book for women features clear linework and engaging details for colored pencils, crayons, gel pens and other favorite coloring supplies.
- A COMFORTING CREATIVE GIFT: A charming choice for women, and adults who enjoy cute animal coloring books, for screen-free relaxation.
- Top-Spiral Lay-Flat Design: The convenient top binding allows the coloring book to rest flat while open, making pages easier to turn and more comfortable to color for both right- and left-handed users.
Split the data before fitting preprocessing
Keep a test set aside before fitting any transformation that learns from the data, including imputers, scalers, and the classifier. The pipeline will learn those steps from the training partition; when asked to predict, it applies the fitted transformations to the test partition without learning from it.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X[["age", "fare", "embarked", "sex", "pclass"]],
y,
test_size=0.2,
random_state=42,
stratify=y,
)
The 20% test fraction and random seed above are choices for a reproducible example, not values established by the scikit-learn documentation as universally best. Stratification is useful for classification when you want the partitions to retain similar class proportions; for a rigorous comparison or hyperparameter search, use an appropriate validation design rather than repeatedly selecting choices based on the test set.
Rank #2
- Ideal Gift: This journal with vibrant embossed patterns makes a thoughtful and versatile gift for occasions like Christmas, birthdays, and more. Convey your best wishes with a present that's both stylish and functional.
- Exquisite Design: Featuring a unique appearance and soft texture, this journal is easy to carry and perfect for use at home, the office, on outdoor adventures, or while traveling. Its classic cover offers excellent protection, while the included strap ensures the contents remain securely organized.
- Perfect Size: Measuring 7.8" × 5" (20 cm × 12.5 cm) with 70 sheets (140 pages), this compact journal is ideal for carrying and writing wherever you go. Easily slip it into your pocket, backpack, or purse for convenient travel. Its versatile design makes it suitable for bullet journaling, daily planning, logging, food tracking, or artistic pursuits like sketching and painting.
- Multifunctional Features: Designed for effortless reading and note-taking, this journal enhances your daily routines, journeys, and work. It includes card slot compartments for organizing essentials like cards, tickets, and photos, along with a zippered page-size slot for securely storing cash, your cell phone, and more.
- Wonderful Gift Idea: Delight your friends, family, and colleagues with this charming and practical journal. It's sure to be appreciated and cherished!
Preprocess numeric and categorical columns separately
Numeric and categorical values are not interchangeable inputs. Missing values need a defined treatment, and category labels such as embarkation port or sex need encoding before a typical classifier can use them. ColumnTransformer routes named columns to separate transformations, while nested preprocessing pipelines keep the operations for each type together.
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
numeric_features = ["age", "fare"]
categorical_features = ["embarked", "sex", "pclass"]
numeric_preprocessor = Pipeline(
steps=[
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler()),
]
)
categorical_preprocessor = Pipeline(
steps=[
("imputer", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore")),
]
)
preprocessor = ColumnTransformer(
transformers=[
("numeric", numeric_preprocessor, numeric_features),
("categorical", categorical_preprocessor, categorical_features),
]
)
The numeric branch fills missing values with the training data’s median and scales the resulting values. The categorical branch fills missing entries with the most frequent training value and creates indicator columns through one-hot encoding. handle_unknown="ignore" allows prediction to proceed if a category appears at transform time that was not seen when the encoder was fitted; the unseen category does not gain its own learned indicator column.
Rank #3
These are reasonable illustrative choices, not a universally optimal recipe. Scaling is often useful for models sensitive to feature scale, but not every classifier needs it. Imputation strategy, category handling, and feature selection should suit the estimator and data. The official example shows column-specific preprocessing and an integrated predictive pipeline; its scikit-learn 1.6.1 version documents the same general pattern.
Join preprocessing and the classifier in one Pipeline
Put the column transformer before a classifier in a top-level Pipeline. This makes fitting and prediction operate through the same ordered workflow, rather than requiring separate manual transformation calls.
from sklearn.linear_model import LogisticRegression
model = Pipeline(
steps=[
("preprocessor", preprocessor),
("classifier", LogisticRegression(max_iter=1000)),
]
)
model.fit(X_train, y_train)
predictions = model.predict(X_test)
fit fits the imputation, scaling, encoding, and classifier stages using training data. predict applies the fitted preprocessing and then asks the classifier for predictions. Keeping these steps together reduces the risk of applying different preprocessing at training and prediction time, and makes the full workflow available to scikit-learn model-selection tools.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallEvaluate predictions without assuming a particular score
A score is meaningful only alongside its metric, split, features, and estimator. The official example is an implementation demonstration, not a statistical report establishing a score that this code should achieve. Calculate the metric you care about on the held-out data rather than assuming a result:
from sklearn.metrics import accuracy_score, classification_report
print("Accuracy:", accuracy_score(y_test, predictions))
print(classification_report(y_test, predictions))
Accuracy is the share of predictions that match the labels, but it may not capture the costs of different mistakes. A classification report also exposes precision, recall, and F1 by class. Choose evaluation measures based on the intended use; do not use test-set results as a repeated tuning signal.
Tune preprocessing and classifier settings together
Because preprocessing and classification are named pipeline steps, search tools can tune parameters from either stage as part of the same estimator. A parameter name uses the step name, two underscores, and the parameter name. For example, the logistic-regression regularization setting is addressed as classifier__C.
from sklearn.model_selection import GridSearchCV, StratifiedKFold
search = GridSearchCV(
estimator=model,
param_grid={
"classifier__C": [0.1, 1.0, 10.0],
"preprocessor__numeric__imputer__strategy": ["mean", "median"],
},
cv=StratifiedKFold(n_splits=5, shuffle=True, random_state=42),
scoring="accuracy",
)
search.fit(X_train, y_train)
best_model = search.best_estimator_
test_predictions = best_model.predict(X_test)
This example searches two regularization values categories and imputation choices; the grid, metric, and fold design are illustrative rather than a performance recommendation. Each cross-validation fit learns preprocessing only from that fold’s training portion because the transformations are inside the pipeline. After selection, assess the selected estimator on the held-out test set once for a less biased final check. Scikit-learn’s official example demonstrates searching parameters across the combined workflow.
Free tools Windows power users keep installed
One-click scans. No signup required.
Optional: request pandas output from transformations
The pipeline does not need pandas-formatted transformed output to fit or predict. If inspecting intermediate transformed data as a DataFrame is useful, scikit-learn also provides an output configuration:
from sklearn import set_config
set_config(transform_output="pandas")
This is an output-format convenience, not a required pipeline component. The official set_output example illustrates the API with Titanic data. Check the documentation for the scikit-learn version installed in your environment, since stable documentation can evolve.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




