// ARTICLE

Predicting Jakarta Travel Time with 372 Features and One Linear Model

A single OLS model turns Jakarta trip context into 372 features, removes 157 extreme residual cases, and produces reproducible travel-time predictions.

Predicting travel time across Jakarta does not require a stack of models to become expressive. One ordinary least-squares regression can represent route, time, weekday, traffic density, and delay patterns when those relationships are encoded before fitting. In this pipeline, 12 raw input columns become a 372-column design matrix, 157 extreme residual cases are removed, and a single coefficient vector produces 3,000 finite travel-time predictions.

The important idea is not the feature count by itself. Each block has a defined job: route-time baselines capture recurring context, density contrasts preserve categorical differences, local linear splines bend the delay response, route interactions allow different slopes, and high-delay steps isolate abrupt changes. The estimator remains linear throughout.

Flow diagram showing trip context becoming five feature blocks with 372 columns, one beta coefficient vector, and 3,000 travel-time predictions

The expressive power comes from the 372-column representation; fitting and prediction still use one coefficient vector.

The expressive power comes from the 372-column representation; fitting and prediction still use one coefficient vector.

Why one linear model can still be expressive

A linear model is linear in its learned coefficients, not necessarily in every raw input. With a feature map phi(x) and coefficient vector beta, the prediction remains y_hat = phi(x) beta. The feature map may convert categories into indicators, create predetermined interactions, or place a continuous value into a piecewise-linear basis. As long as fitting estimates one coefficient vector, the result is still one linear regression.

The final fit uses numpy.linalg.lstsq, which solves a least-squares problem by minimizing the residual norm. There is no decision tree, neural network, boosting stage, ensemble, or second learned model after ordinary least squares. Even the final affine calibration is folded back into the same coefficient vector.

This distinction matters because “simple model” and “simple representation” are different claims. A linear estimator can model rich conditional behavior when the design matrix exposes that behavior explicitly. The tradeoff is equally clear: the model inherits every assumption built into the feature map and cannot improvise when a new category falls outside the trained schema.

Start with a schema that fails loudly

The training data contains 40,000 rows and 13 columns, including the target. The prediction data contains 3,000 rows and 12 input columns. All input names and their order are checked before feature construction so that a renamed, missing, or reordered column causes an explicit failure instead of silently shifting values into the wrong feature.

Only seven inputs enter the final feature map: origin, destination, time, weekday, density at the origin, density at the destination, and delay. The remaining columns are not declared useless in general. They simply do not participate in this particular design. That boundary prevents a local modeling choice from becoming an unsupported universal conclusion.

Missing density values receive a dedicated tuple marker that cannot collide with a real category. Missingness therefore remains a separate state rather than being imputed as an apparently ordinary traffic condition. The approach preserves information, but it also expands the categorical space and depends on enough observations being available for each relevant route-density combination.

The encoder fails closed when prediction data contains a route, route-time-weekday cell, or route-specific density level that never appeared during fitting. That behavior protects column identity and prevents an unseen category from being assigned to an unrelated coefficient. It is suitable for a fixed batch with known coverage, but less suitable for a continuously changing service unless retraining or an explicit fallback policy is added.

Turn each trip into 372 deterministic features

The data contains 10 observed directed routes, four time values, and seven weekdays. Their observed route-time-weekday combinations form 280 contexts. The first design block gives each context its own baseline through one intercept and indicators for every non-reference level.

Five blocks make up the complete matrix:

Five feature blocks forming a deterministic 372-column design matrix
Feature blockColumnsPurpose
Route-time-weekday baseline280Captures the typical travel time of each observed context, including the intercept
Route-specific density contrasts50Allows density categories to have different effects on different routes
Local linear delay spline23Represents a piecewise-linear response across knots from 0.750 to 1.325
Route-specific delay slopes9Allows non-reference routes to respond differently to delay
Route-specific high-delay steps10Adds a correction after delay exceeds 1.25
Total372One deterministic design matrix for one least-squares fit

The total is deliberately bounded: each block answers a different modeling question. Recurring route and calendar structure goes into the baselines, while density contrasts preserve categorical traffic patterns. The spline adds local flexibility, route slopes capture different delay sensitivity, and high-delay steps test whether the upper tail needs a discrete correction.

Reference levels keep the matrix identifiable. One category from a one-hot group is omitted, and one spline basis acts as the reference. The resulting coefficients describe differences relative to those reference terms rather than duplicating the same information across perfectly dependent columns.

Keep nonlinear transformations outside the estimator

Delay is not forced into one global straight line. It is projected onto local linear “hat” functions with knots from 0.750 to 1.325 at intervals of 0.025. Between neighboring knots, weight transfers linearly from one basis column to the next. Omitting one reference basis leaves 23 spline columns.

The transformation is nonlinear with respect to raw delay because the active columns depend on the observation’s position between knots. The estimator is still linear because every basis value is calculated before fitting. Ordinary least squares only chooses the weight assigned to each prepared column.

Route interactions follow the same rule. Nine non-reference routes receive an extra delay slope, while all 10 routes receive an indicator for delay > 1.25. No learned coefficient multiplies another learned coefficient. Every interaction is a fixed number inside the design matrix before the least-squares solver begins.

This structure occupies a middle ground. It captures bends and conditional effects without abandoning a single auditable coefficient vector. Rigidity is the cost: knot locations, thresholds, and interactions are human-defined hypotheses that cannot rescue a weak premise.

Remove only extreme positive residuals

An initial fit produces residuals defined as actual - predicted. Iterative cleaning then identifies unusually large positive residuals with a median absolute deviation rule using the scale factor 1.4826022185. The negative side is not trimmed by the same rule. The statistical building block is documented in SciPy's median absolute deviation reference.

Across five passes, exclusions stabilize at 157: the sequence is 150, 155, 156, 157, then 157 again. The final mask retains 39,843 of 40,000 observations, so cleaning removes only 0.3925%.

One-sided cleaning can be reasonable when extremely slow trips reflect rare incidents, unstable records, or disruptions that a fixed linear specification cannot learn reliably. It also creates a serious limitation. A severe traffic event may contain the most operationally important signal. Removing such cases can improve fit on ordinary conditions, yet weaken the model precisely when traffic becomes worst.

The cleaning mask must therefore be treated as part of the model, not as harmless housekeeping. Any comparison that creates the mask before cross-validation may leak information across folds. The paired comparison below freezes the same mask for both feature sets, so it isolates the incremental feature change; it does not provide a fully unbiased estimate for future traffic. The broader distinction between training decisions and held-out evaluation is covered in scikit-learn's guidance on data leakage.

Test high-delay features as a paired change

The core design contains 362 columns. The complete design adds 10 route-specific high-delay indicators for a total of 372. Both versions use the same retained rows and the same five folds, which makes the difference attributable to the added block rather than to a changed split.

Five-fold delta plot showing four small R-squared improvements and one very small decline after adding route-specific high-delay features

Four folds improve, one declines slightly, and the pooled gain is approximately 0.000102 R-squared.

Four folds improve, one declines slightly, and the pooled gain is approximately 0.000102 R-squared.
Paired R-squared comparison of the 362-column core and 372-column complete design across five folds
FoldCore R²Complete design R²ΔR²
10.9554930.955501+0.000008
20.9551710.955516+0.000344
30.9544770.954533+0.000055
40.9542370.954233−0.000005
50.9554320.955536+0.000104
Pooled0.9549650.955067+0.000102

The result supports a narrow decision. Ten extra columns produce a small pooled improvement, four folds move upward, and one moves slightly downward. The gain is too small to justify a broad claim about high-delay thresholds. It only shows that this particular block adds a modest signal under a controlled paired comparison.

R² measures improvement over a baseline that always predicts the mean target. It equals one minus the residual sum of squares divided by the target’s total variation, and it can be negative when predictions are worse than that baseline. The exact definition and edge cases are available in the scikit-learn R² documentation.

Fold calibration back into the coefficient vector

After the cleaning mask is fixed, the final least-squares fit learns 372 coefficients from 39,843 rows. The matrix has rank 372 and a condition number of approximately 779.18. Full rank means no exact linear dependency remains among the columns on this dataset. A finite condition number suggests a more manageable numerical system than a nearly singular matrix, but it does not guarantee generalization or causal interpretation.

Raw predictions receive an affine calibration of 1.103 * raw + 1.142. That adjustment does not need to remain a separate prediction stage. Every non-intercept coefficient is multiplied by 1.103; the intercept receives the same multiplier plus an added 1.142. Prediction therefore remains a single operation: X_prediction @ beta_final.

The resulting vector contains 3,000 finite values. Its summary is mean=34.6593, standard deviation=16.0017, minimum=13.3856, and maximum=90.8989. Row order is preserved as part of the output contract, and the prediction file contains one numeric column.

A clean rerun from an isolated copy produces the same 3,000-row file with SHA-256 c9135f6ce3978e9809ee4afd8d026becdbcf2d20a1080b910e68d2dc1ea4fdc4. The hash does not prove predictive quality. It proves the narrower property that the pipeline can reproduce the same bytes from the same inputs and code.

Strengths, limits, and a reusable blueprint

Practical strengths and limitations of the linear travel-time pipeline
AreaStrengthLimitation
Trip context280 baselines represent route, time, and weekday combinations explicitlyNew combinations have no safe encoding without a schema update
Delay responseLocal splines and route slopes add flexibility without changing the estimatorKnots, the 1.25 threshold, and basis shape remain design assumptions
Traffic densityMissing values and route-specific categories stay distinguishableSparse interactions require sufficient coverage and careful reference levels
CleaningOnly 0.3925% of rows in the extreme positive-residual tail are removedValid severe congestion may be discarded with anomalous records
Numerical behaviorThe final matrix is full rank with a condition number near 779Numerical stability does not prove transfer to new traffic patterns
ValidationCore and complete designs share identical rows and foldsA precomputed clean mask limits the strength of out-of-sample claims
OperationsOne solver, one coefficient vector, one matrix multiplication, and a reproducible output fileFail-closed category handling requires retraining or a fallback for new routes

The reusable pattern starts with a clear prediction contract and a schema that rejects drift. Recurring context becomes explicit baselines; only interactions tied to a concrete hypothesis are added. Spline bases and thresholds remain fixed transformations, not extra estimators. Each new block is compared on identical folds. When an unbiased deployment estimate matters, cleaning belongs inside the validation design. Affine calibration can then be folded into the final coefficient vector, followed by a deterministic output check.

The 372 columns are useful because their roles are inspectable, not because a larger matrix is inherently better. Feature engineering gives a linear model a richer vocabulary, while strict schema checks, paired validation, and reproducible output keep that vocabulary accountable. The next practical step is to run the same paired-ablation pattern against one new feature block, then retain it only when the improvement survives identical folds.

Read next

Articles sharing this post's topics, followed by the latest entries.

View all articles
// READER SIGNAL

Reader notes

0 comments
// FIELD NOTES

Comments

Checking your session...

Loading reader comments...

Open sourceBack to all articles
Jakarta Travel Time Prediction with Linear Regression · Mukhtada