Store Expansion Cost Prediction using Stepwise Regression in ML

Get Ready for Your Dream Job: Click, Learn, Succeed, Start Now!

Retail chains planning to expand must estimate the capital required to open new outlets—incorporating site acquisition, fit‑out, staffing, and inventory costs. In this project, we will predict the total expansion cost for a prospective store using location attributes (population density, median income, competitor count), store format (size, format type), and historical performance of nearby outlets. By applying stepwise regression, we’ll identify the strongest cost drivers and develop a lean, interpretable linear model that helps finance teams budget expansions more accurately.

Libraries Required

import pandas as pd               # Data manipulation  
import numpy as np                # Numerical operations  
import statsmodels.api as sm      # Statistical modeling (OLS)  
from sklearn.model_selection import train_test_split   # Data splitting  
from sklearn.metrics import r2_score, mean_squared_error  # Evaluation  
import matplotlib.pyplot as plt   # Visualization  

Dataset

Locating New Stores

Step-by-Step Code Implementation

Data Loading & Initial Inspection

We import a store‐location dataset containing demographic and competitive features, along with a proxy cost column (Expansion_Cost). Initial inspection (.info(), .describe()) checks for data types, distributions, and completeness.

# Block 1: Load dataset
# We use the "Locating New Stores" dataset as a proxy, which includes site demographics and nearby store performance :contentReference[oaicite:0]{index=0}  
df = pd.read_csv("locating_new_stores.csv")

# Inspect data
print(df.head())
print(df.info())
print(df.describe())

Data Preprocessing

The categorical Store_Format is one‑hot encoded; missing rows are dropped. Predictors (X) include numeric site attributes and encoded formats; the response (y) is Expansion_Cost. We split the data 80/20 for training and testing.

# Block 2: Encode categoricals & handle missing values
df_enc = pd.get_dummies(df, 
                        columns=["Store_Format"],  # e.g., 'Small', 'Standard', 'Large'
                        drop_first=True)

# Assume the dataset has columns:
# 'Population_Density', 'Median_Income', 'Competitor_Count', 
# 'Avg_Nearby_Sales', 'Distance_to_HQ', and our target 'Expansion_Cost'
df_enc = df_enc.dropna()

# Define predictors and target
X = df_enc.drop("Expansion_Cost", axis=1)
y = df_enc["Expansion_Cost"]

# Split into train/test sets
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

Stepwise Regression Function

Our stepwise_selection function alternates forward inclusion—adding the excluded predictor with the smallest p‑value below 0.01—and backward elimination—removing the included predictor with the largest p‑value above 0.05—until convergence, yielding a succinct set of cost drivers.

# Block 3: Forward–backward stepwise selection
def stepwise_selection(X, y,
                       initial_list=[],
                       threshold_in=0.01,
                       threshold_out=0.05,
                       verbose=True):
    included = list(initial_list)
    while True:
        changed = False

        # Forward step: add best excluded variable
        excluded = list(set(X.columns) - set(included))
        pvals = pd.Series(index=excluded, dtype=float)
        for col in excluded:
            model = sm.OLS(y, sm.add_constant(X[included + [col]])).fit()
            pvals[col] = model.pvalues[col]
        best_pval = pvals.min()
        if best_pval < threshold_in:
            best_var = pvals.idxmin()
            included.append(best_var)
            changed = True
            if verbose:
                print(f"Add  {best_var:30} p-value {best_pval:.6f}")

        # Backward step: remove worst included variable
        model = sm.OLS(y, sm.add_constant(X[included])).fit()
        pvals_incl = model.pvalues.iloc[1:]  # exclude intercept
        worst_pval = pvals_incl.max()
        if worst_pval > threshold_out:
            worst_var = pvals_incl.idxmax()
            included.remove(worst_var)
            changed = True
            if verbose:
                print(f"Drop {worst_var:30} p-value {worst_pval:.6f}")

        if not changed:
            break

    return included

Model Building & Evaluation

Using the selected features, we fit an Ordinary Least Squares regression via statsmodels. The .summary() output provides coefficient estimates, standard errors, p‑values, R², and diagnostic statistics, offering insight into each predictor’s effect.

We evaluate on the held‑out test set and compute R² (explained variance) and RMSE (root‑mean‑square error) to quantify generalisation performance.

# Block 4: Feature selection
selected_features = stepwise_selection(X_train, y_train)

# Fit final OLS model
X_train_sel = sm.add_constant(X_train[selected_features])
model = sm.OLS(y_train, X_train_sel).fit()
print(model.summary())

# Predict on test set
X_test_sel = sm.add_constant(X_test[selected_features])
y_pred = model.predict(X_test_sel)

# Compute performance metrics
print("Test R²:", r2_score(y_test, y_pred))
print("Test RMSE:", np.sqrt(mean_squared_error(y_test, y_pred)))

Residual Diagnostics

A scatter plot of residuals versus predicted costs checks for non‑random patterns or heteroscedasticity, validating key linear regression assumptions.

# Block 5: Residual plot
residuals = y_test - y_pred
plt.scatter(y_pred, residuals)
plt.axhline(0, linestyle="--")
plt.xlabel("Predicted Expansion Cost")
plt.ylabel("Residuals")
plt.title("Residuals vs. Predicted Cost")
plt.show()

Summary

By applying stepwise regression to site‐location data, we isolate the principal drivers of store expansion cost—such as local population density, median income, competitor density, and store format—while pruning less informative factors. The resulting linear model balances interpretability with accuracy (high test‐set R², low RMSE), equipping retail finance teams with a transparent forecasting tool to budget new store rollouts and optimise capital allocations.

We work very hard to provide you quality material
Could you take 15 seconds and share your happy experience on Google | Facebook


PythonGeeks Team

The PythonGeeks Team delivers expert-driven tutorials on Python programming, machine learning, Data Science, and AI. We simplify Python concepts for beginners and professionals to help you master coding and advance your career.

Leave a Reply

Your email address will not be published. Required fields are marked *