Guide
Feature Scaling: Why It Matters and What Happens If You Do It Wrong
Yasin Polat
the AI · Guide
Why Feature Scaling is Vital and What Happens if It's Applied Incorrectly?
In machine learning, everything starts with data. However, for this data to be truly “understandable” by the model, it is not enough to simply collect it; it must be prepared correctly. This is because raw data can be on different scales, in different units, or have very wide ranges. This is where feature scaling comes into play.
Feature scaling ensures that the model treats all features fairly by bringing the numerical values in the dataset to a similar range. Although this process is often overlooked in the machine learning process, it is one of the most critical steps that directly affects the success of the model. Scaling prevents the algorithm from “suppressing” small numerical values with large ones in the data, so that each feature contributes equally to the prediction process.
In other words, feature scaling brings the model into balance. By bringing the data to the same level, it allows the model to decide which feature is truly important based on the information it contains, not on differences in measurement. This helps the model learn faster and produce more accurate results.
Let's Start with a Simple Example
Let's say we want to build a house price prediction model. We have the following two features:
- Square meters: 60 – 300
- Number of rooms: 1 – 6
If we feed this data into the model in its raw form, the square footage values (60, 120, 250, etc.) will be numerically much larger than the number of rooms. Mathematically, the model will pay more “attention” to these large numbers. This means that when predicting the house price, the model will overemphasize square footage and not sufficiently consider the impact of the number of rooms. However, the number of rooms is also an important factor affecting the price.
Due to this imbalance in the data, the model may unnecessarily prioritize certain features, leading to inaccurate or biased predictions.
Now let's think about it: If we bring square footage and number of rooms to the same scale (for example, by converting them to values between 0 and 1), the model will now give equal weight to both features. This ensures that both square footage and number of rooms are fairly evaluated in the price prediction process.
Scaling makes the model's decision-making mechanism more balanced, its predictions more reliable, and its results more understandable.
The importance of feature scaling is not limited to this example. In the real world, the ranges and units of features vary greatly across many datasets. Therefore, feeding data directly into the model without scaling it appropriately can disrupt the model's learning process; it may unnecessarily emphasize some features or cause it to completely ignore others.
1. Why Do We Perform Feature Scaling?
Using the KNN (K-Nearest Neighbor) Example
To understand the importance of feature scaling, it is necessary to know how algorithms use data. Especially distance-based algorithms like KNN or K-Means make decisions by measuring similarities or distances between data points. If features are on different scales, some become dominant over others and the model is misguided.
A Simple Bioinformatics-Focused Scenario
Let's say we have a genomic dataset and our goal is to predict disease risk. Let's say we have two features:
-
Gene A expression level: 0 – 1000
-
Gene B expression level: 0 – 5
What happens when we feed this raw data into the KNN model?
-
Since the numerical value of Gene A is very large, Gene A takes all the weight in the distance calculation.
-
The effect of Gene B is almost negligible.
-
The contribution of a gene that could be biologically important is not reflected in the model.
If we feed this data into the KNN model in its raw form, the high numerical values of Gene A completely overshadow the contribution of Gene B. In other words, the algorithm almost ignores the role of Gene B in disease risk. This situation leads to the loss of the effect of a gene that could be biologically significant. However, when we scale the data, for example, by compressing the expression levels of both genes between 0 and 1, the model now considers the effects of both genes equally. This allows the algorithm to correctly learn biologically meaningful relationships, and the predictions become more reliable.
After Feature Scaling
If we scale both gene values to the 0–1 range, for example:
-
KNN assigns equal weight to both genes.
-
The biological importance of Gene B is now correctly reflected in the model.
-
The algorithm learns relationships more accurately, and the results become more reliable.
Therefore, feature scaling is a critical step, especially in distance-based methods.
import numpy as np
import pandas as pd
from sklearn.preprocessing import MinMaxScaler, StandardScaler
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
# Sample genomic data
# Gene A expression: 0-1000 range
# Gene B expression: 0-5 range
data = {
'GenA': [100, 500, 800, 200, 900, 50, 700, 400, 600, 300],
'GenB': [1, 4, 3, 2, 5, 0, 4, 2, 3, 1],
'Disease': [0, 1, 1, 0, 1, 0, 1, 0, 1, 0] # target variable
}
df = pd.DataFrame(data)
# Features and target
X = df[['GenA', 'GenB']]
y = df['Disease']
# Split the data into training and test sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# KNN raw data
knn = KNeighborsClassifier(n_neighbors=3)
knn.fit(X_train, y_train)
pred_original = knn.predict(X_test)
acc_original = accuracy_score(y_test, pred_original)
print(f"KNN accuracy with raw data: {acc_original:.2f}")
# Min-Max Normalization
scaler = MinMaxScaler()
X_train_norm = scaler.fit_transform(X_train)
X_test_norm = scaler.transform(X_test)
knn.fit(X_train_norm, y_train)
pred_norm = knn.predict(X_test_norm)
acc_norm = accuracy_score(y_test, pred_norm)
print(f"KNN accuracy with normalized data: {acc_norm:.2f}")
# Standardization
scaler_std = StandardScaler()
X_train_std = scaler_std.fit_transform(X_train)
X_test_std = scaler_std.transform(X_test)
knn.fit(X_train_std, y_train)
pred_std = knn.predict(X_test_std)
acc_std = accuracy_score(y_test, pred_std)
print(f"KNN accuracy with standardized data: {acc_std:.2f}")
Here, the ranges of Gene A and Gene B are very different. With raw data, KNN places too much importance on Gene A. When Min-Max normalization or standardization is applied, KNN now evaluates both genes with equal importance. The code's output will typically show lower accuracy with raw data and higher accuracy after scaling.
Gradient Descent
Algorithms such as linear regression, logistic regression, and artificial neural networks learn using the gradient descent method. In other words, the model is continuously updated step by step to minimize the error function. At each step, the model slightly adjusts its parameters based on the slope of the error function.
But there is a critical point here: If the data is on different scales, the step sizes of gradient descent will also be different. For example, one feature may take values between 1 and 5, while another feature may take values between 100 and 10,000. In this case, the gradient descent algorithm:
-
Takes very large steps for the large interval feature,
-
Takes very small steps for the small interval feature.
-
As a result, the model progresses toward the optimum point in an unstable and erratic manner.
This situation:
-
Slows down the learning process,
-
The model may not reach the optimum,
-
The error function fluctuates and instability is observed during training.
When we scale the data, this problem disappears. Since all features are on the same scale, gradient descent takes balanced steps, and the learning process becomes more stable and faster. Let's illustrate this situation using a simple linear regression example.
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import StandardScaler
# Sample data: two features, different scales
# Feature 1: 1-5, Feature 2: 100-1000
X = np.array([[1, 100],
[2, 200],
[3, 300],
[4, 400],
[5, 500]])
y = np.array([10, 20, 30, 40, 50])
# Raw data model
model_raw = LinearRegression()
model_raw.fit(X, y)
print("Raw data with coefficients:", model_raw.coef_)
# Let's standardize the features
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
model_scaled = LinearRegression()
model_scaled.fit(X_scaled, y)
print("Coefficients with standardized data:", model_scaled.coef_)
Raw data and coefficients are more dependent on large-scale features. In other words, the model almost ignores small-scale features. With standardized data, since every feature is on the same scale, the model treats both features fairly. This simple example demonstrates the importance of balancing the steps of gradient descent. In real large datasets, this situation directly affects the model's speed and stability.
More consistent numerical behavior
Some machine learning algorithms are highly sensitive to the scales of the data mathematically. For example:
- SVM (Support Vector Machines),
- PCA (Principal Component Analysis),
- Regularized regression models such as Lasso or Ridge Regression.
These algorithms perform mathematical calculations on the data, such as standard deviation, variance, or covariance. If the features are on different scales, these calculations become biased. For example, one feature may take values between 0 and 1, while another feature may take values between 0 and 10,000. In this case, the large-scale feature becomes much more dominant than the small-scale feature. The mathematical structure of the model is optimized according to this dominant feature, and the small-scale feature is almost ignored. This situation becomes even more critical in dimension reduction techniques such as PCA.
PCA finds the directions that maximize the variance in the data. If the data is unscaled, high-range features dominate the variance, and the contribution of low-range features is almost negligible. As a result, the model cannot learn the true structure of the data and prioritizes incorrect directions during dimension reduction. Similarly, in the SVM algorithm, kernel functions and distance calculations are sensitive to the scale of the data.
Scale differences make it difficult for the model to find the correct classification boundaries. In Lasso and Ridge regression, penalty terms create different effects on features with different scales. A large-scale feature is more affected by the regularization (penalty) term, while a small-scale feature is less affected. This can disrupt the model's learning process and lead to incorrect coefficient estimates.
What Methods Can We Use?
Feature scaling is a fundamental preprocessing step used to bring data to the same scale so that the model can evaluate all features fairly. The two most common methods in this process are Normalization (Min-Max Scaling) and Standardization (Z-score Scaling).
Normalization (Min-Max Scaling)
Normalization compresses data into a specific range (usually between 0 and 1). This ensures that all values are represented on the same scale and large numerical differences do not dominate the model's learning process.
Min-Max Normalization Formula
Normalization (scaled value $x'$):
$$x' = \frac{x - x_{min}}{x_{max} - x_{min}}$$
Example Application
If a student's exam score is $x=70$, the lowest score is $x_{min}=50$, and the highest score is $x_{max}=100$, then:
$$x' = \frac{70 - 50}{100 - 50} = \frac{20}{50} = 0.4$$
This student's score is now represented by a value between $0$ and $1$ ($0.4$).
When is it used?
Normalization is preferred especially in the following situations:
-
When the distribution of the data is unknown: If you do not know whether your data follows a normal distribution and you only want to compress the values into a specific range, normalization is an appropriate method.
-
When using distance-based algorithms: Algorithms such as KNN, K-Means, or Neural Networks use the distances between data points. If features are on different scales, features with large values dominate the distance, and features with small values are neglected. Normalization eliminates this bias.
-
If values need to be kept within a certain range: Some algorithms or applications require input values to be within a specific range. For example, some artificial neural network layers typically expect data in the range 0–1 or -1–1. In this case, using normalization is unavoidable.
Standardization (Z-score Scaling)
Standardization is the process of transforming data so that its mean is 0 and its standard deviation is 1. In other words, each data point expresses how far it is from the mean of the relevant feature in terms of standard deviation units. This method is particularly effective when the data distribution is normal or approximately normal.
Standardization (Z-Score Normalization) Formula
Standardization transforms a data point ($\mathbf{x}$) into a distribution with mean $\mathbf{0}$ and standard deviation $\mathbf{1}$.
Standardized value ($\mathbf{x'}$):
$$x' = \frac{x - \mu}{\sigma}$$
-
$\mathbf{x}$: Original value
-
$\mathbf{\mu}$: Mean of the feature
-
$\mathbf{\sigma}$: Standard deviation of the feature
-
$\mathbf{x'}$: Standardized value (Z-Score)
Example Application
If the mean $\mathbf{\mu=50}$ and standard deviation $\mathbf{\sigma=10}$ of a characteristic are given, the new value of $\mathbf{x=70}$ is:
$$x' = \frac{70 - 50}{10} = \frac{20}{10} = 2$$
This result ($\mathbf{x'=2}$) indicates that the value $\mathbf{70}$ is located 2 standard deviations above the mean.
When is Standardization (Z-Score) Used?
Standardization is particularly advantageous and preferred over Min-Max Normalization in the following situations:
1. When the Data Distribution is Approximately Normal
-
Standardization is more compatible with the assumption of normal distribution.
-
Even if features have different ranges, the data is shifted to a central location ($\mu=0$) and evaluated on the same scale ($\sigma=1$).
2. If Scale-Sensitive Algorithms Are Used
-
This is particularly necessary for algorithms based on distance or covariance calculations:
-
SVM (Support Vector Machines): Kernel functions are scale-sensitive.
-
PCA (Principal Component Analysis): Covariance and variance calculations are affected by scale.
-
Logistic Regression: Estimating regression coefficients can lead to large deviations in unscaled data.
-
-
Standardization ensures that these algorithms work balanced and accurate mathematically.
3. If You Want to Reduce the Effect of Outliers
-
Min-Max Normalization compresses extreme values into the $0-1$ range, but the effect of extreme values is very pronounced and can distort the entire data set.
-
Standardization expresses outliers in terms of standard deviation, so extreme values have a more balanced effect when scaled and are not overly compressed.
4. When Fairness Among Features is Desired
-
Especially when there are numerical features of different magnitudes, standardization ensures that each feature is evaluated with equal importance.
-
This allows gradient descent-based models (Linear Regression, Logistic Regression, Neural Networks) to learn more stably and quickly.
What Happens If Feature Scaling Is Done Incorrectly?
Certain mistakes made during feature scaling can seriously affect the performance and generalization ability of machine learning models. These mistakes often carry the risk of data leakage.
1. Scaling the Train and Test Data Separately (Risk of Data Leakage)
This is one of the most common and serious mistakes made in feature scaling.
-
Reason for Error: If you scale the train and test sets separately, the data scale seen by the model changes each time. The scaler uses statistics such as the mean ($\mu$) and standard deviation ($\sigma$) of the data in the test set.
-
For example, the mean expression of a gene in the train set may be 100, while the mean in the test set may be 150. Scaling the test set separately causes the model to use a different scale than the scaling parameters it learned from the train set.
-
Result: The model cannot find the scale it expects on the test data, which leads to a drop in model performance during testing and production. The ability to generalize to real-world data is lost.
Correct Method:
-
Fit the scaler only to the TRAIN data using
fit(). This step ensures that the scaling parameters ($\mu$, $\sigma$, $x_{min}$, $x_{max}$) are learned only from the training data. -
Then use the same scaler to
transform()both the TRAIN and TEST data.- This way, the model evaluates the test data using the fixed scale it learned on the train set and accurately shows its true performance.
2. Scaling Categorical Data (Loss of Meaning)
Some developers scale One-Hot Encoded (OHE) categorical data with the mindset of “let's quantify everything.” This is a mistake.
-
Reason for Error: Categorical data (e.g., $\text{female}=0, \text{male}=1$) mathematically represents separate classes, not a difference in quantity.
-
If we attempt to standardize these $\mathbf{0}$ and $\mathbf{1}$ values, the new values are no longer $\mathbf{0}$ and $\mathbf{1}$ (e.g., $-0.5$ and $0.5$).
-
The model interprets this numerical change as a real quantitative difference, and the accuracy of the categorical information is lost.
-
Correct Method:
-
Never scale categorical data (both OHE-encoded and label-encoded).
-
Only transform continuous and numeric features using Normalization or Standardization.
3. Forced Scaling in the Wrong Model (Unnecessary Processing)
Some machine learning algorithms are inherently insensitive to the scale of the data and do not require scaling.
-
Examples: Tree-based models such as Decision Tree, Random Forest, Gradient Boosting (XGBoost, LightGBM).
-
Why are they insensitive? These models operate on data through split points (e.g., “If age $> 35$, go right”). Split decisions are influenced not by the numerical magnitude of the data, but by whether the value is present at that point.
-
Side Effect: Scaling these models is generally harmless, but it creates unnecessary computational load. This wastes time and resources. Scaling is only mandatory for distance and gradient-based algorithms.
Bültenime Abone Olun
Tüm güncellemeleri doğrudan benden almak için abone ol!

