October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

How Do You Use a Scikit-Learn Pipeline on Titanic Data?

Learn to fetch Titanic data, preprocess numeric and categorical columns with ColumnTransformer, combine the steps in a scikit-learn Pipeline, and evaluate or tune the complete estimator.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A scikit-learn pipeline lets you treat preprocessing and classification as one estimator: fit the complete workflow on training data, then use that same workflow to predict on held-out data or new passengers. For Titanic data, a ColumnTransformer can impute and scale numeric fields while imputing and one-hot encoding categorical fields; a Pipeline then connects those transformations to a classifier.

Load the Titanic data and inspect its columns

The official scikit-learn example fetches the Titanic dataset from OpenML and returns feature data as a pandas DataFrame alongside the target:

As an Amazon Associate I earn from qualifying purchases.

from sklearn.datasets import fetch_openml

X, y = fetch_openml(
    "titanic",
    version=1,
    as_frame=True,
    return_X_y=True,
)

print(X.dtypes)
print(X.isna().sum())
print(y.head())

Here, X contains input features and y is the survival target, named survived in the dataset. Inspect actual column names and missing values before choosing fields: the official example uses age and fare as numeric features, and embarked, sex, and pclass as categorical features. See the official mixed-type Titanic example.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a compact demonstration, use those five fields. This is a feature-selection choice for the example, not a claim that they are the only useful predictors. If you change the fields, make sure the names match the fetched DataFrame.

#1 Best Overall
Titanic Coloring Book for Adults, Womens Book with Top Spiral Binding
  • TITANIC SHIP SCENES, PASSENGERS AND VINTAGE DETAILS: Color historical ocean liner illustrations featuring promenade decks, elegant travelers, mothers and children, photographers, portholes, deck chairs, luggage, ship equipment, cabins, nautical details, and Edwardian maritime scenes created for Titanic fans, history lovers, collectors, seniors, beginners, and adult colorists.
  • Thick Cardstock Paper: Each design is printed on substantial cardstock for a sturdier coloring surface. The single-sided format gives every illustration its own page and helps protect the next design while coloring.
  • Detailed Designs for Adults: This spiral adult coloring book for women features clear linework and engaging details for colored pencils, crayons, gel pens and other favorite coloring supplies.
  • A COMFORTING CREATIVE GIFT: A charming choice for women, and adults who enjoy cute animal coloring books, for screen-free relaxation.
  • Top-Spiral Lay-Flat Design: The convenient top binding allows the coloring book to rest flat while open, making pages easier to turn and more comfortable to color for both right- and left-handed users.

Split the data before fitting preprocessing

Keep a test set aside before fitting any transformation that learns from the data, including imputers, scalers, and the classifier. The pipeline will learn those steps from the training partition; when asked to predict, it applies the fitted transformations to the test partition without learning from it.

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X[["age", "fare", "embarked", "sex", "pclass"]],
    y,
    test_size=0.2,
    random_state=42,
    stratify=y,
)

The 20% test fraction and random seed above are choices for a reproducible example, not values established by the scikit-learn documentation as universally best. Stratification is useful for classification when you want the partitions to retain similar class proportions; for a rigorous comparison or hyperparameter search, use an appropriate validation design rather than repeatedly selecting choices based on the test set.

Rank #2
InnoBeta Titanic Gifts Journal Notebook, Gifts for Titanic Lovers on Christmas, Birthday, Sketchbook, Travel Diary, Lined Planner, Faux Leather, 7.8 * 5 Inches
  • Ideal Gift: This journal with vibrant embossed patterns makes a thoughtful and versatile gift for occasions like Christmas, birthdays, and more. Convey your best wishes with a present that's both stylish and functional.
  • Exquisite Design: Featuring a unique appearance and soft texture, this journal is easy to carry and perfect for use at home, the office, on outdoor adventures, or while traveling. Its classic cover offers excellent protection, while the included strap ensures the contents remain securely organized.
  • Perfect Size: Measuring 7.8" × 5" (20 cm × 12.5 cm) with 70 sheets (140 pages), this compact journal is ideal for carrying and writing wherever you go. Easily slip it into your pocket, backpack, or purse for convenient travel. Its versatile design makes it suitable for bullet journaling, daily planning, logging, food tracking, or artistic pursuits like sketching and painting.
  • Multifunctional Features: Designed for effortless reading and note-taking, this journal enhances your daily routines, journeys, and work. It includes card slot compartments for organizing essentials like cards, tickets, and photos, along with a zippered page-size slot for securely storing cash, your cell phone, and more.
  • Wonderful Gift Idea: Delight your friends, family, and colleagues with this charming and practical journal. It's sure to be appreciated and cherished!

Preprocess numeric and categorical columns separately

Numeric and categorical values are not interchangeable inputs. Missing values need a defined treatment, and category labels such as embarkation port or sex need encoding before a typical classifier can use them. ColumnTransformer routes named columns to separate transformations, while nested preprocessing pipelines keep the operations for each type together.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

numeric_features = ["age", "fare"]
categorical_features = ["embarked", "sex", "pclass"]

numeric_preprocessor = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ]
)

categorical_preprocessor = Pipeline(
    steps=[
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("onehot", OneHotEncoder(handle_unknown="ignore")),
    ]
)

preprocessor = ColumnTransformer(
    transformers=[
        ("numeric", numeric_preprocessor, numeric_features),
        ("categorical", categorical_preprocessor, categorical_features),
    ]
)

The numeric branch fills missing values with the training data’s median and scales the resulting values. The categorical branch fills missing entries with the most frequent training value and creates indicator columns through one-hot encoding. handle_unknown="ignore" allows prediction to proceed if a category appears at transform time that was not seen when the encoder was fitted; the unseen category does not gain its own learned indicator column.

These are reasonable illustrative choices, not a universally optimal recipe. Scaling is often useful for models sensitive to feature scale, but not every classifier needs it. Imputation strategy, category handling, and feature selection should suit the estimator and data. The official example shows column-specific preprocessing and an integrated predictive pipeline; its scikit-learn 1.6.1 version documents the same general pattern.

Join preprocessing and the classifier in one Pipeline

Put the column transformer before a classifier in a top-level Pipeline. This makes fitting and prediction operate through the same ordered workflow, rather than requiring separate manual transformation calls.

from sklearn.linear_model import LogisticRegression

model = Pipeline(
    steps=[
        ("preprocessor", preprocessor),
        ("classifier", LogisticRegression(max_iter=1000)),
    ]
)

model.fit(X_train, y_train)
predictions = model.predict(X_test)

fit fits the imputation, scaling, encoding, and classifier stages using training data. predict applies the fitted preprocessing and then asks the classifier for predictions. Keeping these steps together reduces the risk of applying different preprocessing at training and prediction time, and makes the full workflow available to scikit-learn model-selection tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluate predictions without assuming a particular score

A score is meaningful only alongside its metric, split, features, and estimator. The official example is an implementation demonstration, not a statistical report establishing a score that this code should achieve. Calculate the metric you care about on the held-out data rather than assuming a result:

from sklearn.metrics import accuracy_score, classification_report

print("Accuracy:", accuracy_score(y_test, predictions))
print(classification_report(y_test, predictions))

Accuracy is the share of predictions that match the labels, but it may not capture the costs of different mistakes. A classification report also exposes precision, recall, and F1 by class. Choose evaluation measures based on the intended use; do not use test-set results as a repeated tuning signal.

Tune preprocessing and classifier settings together

Because preprocessing and classification are named pipeline steps, search tools can tune parameters from either stage as part of the same estimator. A parameter name uses the step name, two underscores, and the parameter name. For example, the logistic-regression regularization setting is addressed as classifier__C.

from sklearn.model_selection import GridSearchCV, StratifiedKFold

search = GridSearchCV(
    estimator=model,
    param_grid={
        "classifier__C": [0.1, 1.0, 10.0],
        "preprocessor__numeric__imputer__strategy": ["mean", "median"],
    },
    cv=StratifiedKFold(n_splits=5, shuffle=True, random_state=42),
    scoring="accuracy",
)

search.fit(X_train, y_train)
best_model = search.best_estimator_
test_predictions = best_model.predict(X_test)

This example searches two regularization values categories and imputation choices; the grid, metric, and fold design are illustrative rather than a performance recommendation. Each cross-validation fit learns preprocessing only from that fold’s training portion because the transformations are inside the pipeline. After selection, assess the selected estimator on the held-out test set once for a less biased final check. Scikit-learn’s official example demonstrates searching parameters across the combined workflow.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Optional: request pandas output from transformations

The pipeline does not need pandas-formatted transformed output to fit or predict. If inspecting intermediate transformed data as a DataFrame is useful, scikit-learn also provides an output configuration:

from sklearn import set_config

set_config(transform_output="pandas")

This is an output-format convenience, not a required pipeline component. The official set_output example illustrates the API with Titanic data. Check the documentation for the scikit-learn version installed in your environment, since stable documentation can evolve.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.