Classification & Regression

Predictive classification estimates which category a case belongs to, while predictive regression estimates a numerical value. These models can prioritise large volumes of work, but an overall accuracy score can conceal costly errors or weak performance for important groups. Marketways defines the action each prediction will trigger and evaluates mistakes by their business consequence. The model is judged by whether it improves the operating decision, not by an impressive score in isolation.

The decision this method supports

We use Classification & Regression to help clients answer: Can the available features predict the required outcome reliably enough for the intended use?

How the method works

Supervised machine learning learns from examples where the required outcome is already known. Classification predicts a category, while regression predicts a numerical value. Performance must be tested on unseen observations that resemble the intended operating environment.

A business example

A service team may classify incoming requests by the team best placed to respond. Evaluation should examine errors by request type as well as overall accuracy, because misrouting an urgent safety issue has a different consequence from misrouting a routine enquiry.

How the client uses the result

Predictive classification and regression help a business anticipate which category or value is likely for a new case. The benefit is faster prioritisation, routing or intervention across volumes of work that people could not assess consistently one case at a time.

What we deliver

We produce predictions, scores, groups, alerts or extracted information together with validation evidence. A manager needs operating thresholds, error consequences, escalation rules and a plan for monitoring change after deployment.

Limits and complementary methods

Prediction is designed to anticipate outcomes, not explain their causes. A highly accurate model may still be unsuitable when errors are costly, unevenly distributed or impossible to act upon.

Selected methods and techniques

We select from these established methods according to the decision, evidence and operating conditions.

  • Binary / multiclass classification: Assign cases to one of two or more predefined outcome categories using labelled examples or an explicit rule. Define each class, eligible population and prediction horizon; ensure predictors would be available when the prediction is made; separate training from later representative evaluation; and address label error, class imbalance and changing populations. Report errors for each class, including false alarms and missed events, and choose any operating threshold from the decision costs and capacity. A class score may rank cases without being a calibrated probability. Classification predicts a category represented in the data; it does not explain why the event occurs or establish that an action will change it.
  • Classification and regression modelling: Use input variables to predict either a categorical outcome through classification or a numerical outcome through regression.
  • Driver analysis / regression: Use regression or a related statistical model to estimate how a defined outcome is associated with measured factors while accounting for the other variables represented in the model. State the outcome, population, period, predictors and comparison; document missing data, influential cases, dependence, nonlinearities and subgroup differences; and report model performance, uncertainty and validation. Results depend on the variables and design available, so cross-sectional or self-reported associations do not establish that changing a predictor will cause the outcome to change. The method is reusable across domains: it may populate an Engagement Driver Model for employees or a Satisfaction / Loyalty Driver Model for customers without making those outputs the same object.
  • Gradient boosting / random forests: Apply tree-ensemble algorithms to a forecasting target using lagged outcomes, calendar fields and other predictors that are legitimately available at the forecast date. Random forests average predictions from many varied trees, while gradient boosting builds trees in sequence to reduce remaining error. Define how each forecast horizon is produced, tune and compare the models using time-ordered validation, check important segments and unusual periods and guard against future information leaking through features or preprocessing. These are specific machine-learning algorithms within Machine-Learning Forecasting. Their variable-importance or contribution measures describe the fitted prediction and do not by themselves establish causal drivers.
  • Linkage / regression to customer outcomes: Link measures of customers' experiences to later outcomes and estimate how they are related, allowing for relevant customer differences. A relationship in the records alone does not establish cause and effect.
  • NLP / text classification for issue logs: Assign issue-log text to predefined operational categories using labelled examples, rules or a supervised language model. Define the record unit, taxonomy and version, single-label or multi-label rule, training and later evaluation samples, treatment of duplicates and missing context, and the routing of low-confidence or novel cases. Report per-category precision, recall and common confusions, including rare high-severity classes, and monitor changes in vocabulary and source systems. Classification applies known labels; issue clustering discovers candidate groupings, and neither a label nor a model explanation proves the root cause.
  • Pay-equity regression analysis: Estimate pay differences between groups using regression models that account for specified pay-related factors, such as job, grade, experience, location and performance, and report the remaining adjusted difference with its uncertainty. The factors included must be justified against the question being asked: controlling for grade may be appropriate for estimating within-grade pay differences but can remove disparities arising from unequal access to higher grades when examining the organisation-wide earnings gap. The method identifies patterns requiring investigation and does not by itself prove discrimination.
  • Regression / dynamic regression: Fit a regression forecast that relates a time-indexed target to observed, planned or separately forecast predictors and, where needed, to lags, trends, seasonal terms or serially related errors. Define the target, units, population, vintage and horizon, make future predictor paths explicit, use only information available at each historical forecast date and test the model on later periods. Dynamic regression is the broader family of regression specifications that represent dependence through time; ARIMAX is one form that models remaining errors with an ARIMA structure. Coefficients can support prediction without showing that changing a predictor would cause the forecast outcome to change.
  • Regression / econometric modelling: Estimate how a clearly defined outcome varies with selected explanatory factors in a stated population and period. Specify the functional form, timing, controls and error structure; inspect missing data, outliers, dependence, instability and model fit; quantify uncertainty; and test performance on later or held-out observations when forecasting. A useful predictive relationship does not show that deliberately changing a factor will change demand. Causal interpretation requires a defensible design and assumptions about confounding, selection and reverse causation. Forecasts outside the observed ranges or after structural change require separate support rather than automatic extrapolation.
  • Regression / multilevel modelling: Relate a defined outcome to one or more measured predictors while reporting the size and uncertainty of the relationships. Multilevel modelling extends ordinary regression when observations are nested or repeated, such as assessments within employees and employees within teams, so person and group variation are represented separately. Specify the unit, period, variables, functional form, missing-data treatment and validation. The coefficients are adjusted associations unless a credible causal design justifies an intervention claim.
  • Regression / multivariate analysis: Use regression to model one defined outcome from one or more predictors, or use a named multivariate method when several outcomes or variables must be analysed jointly. State the analytical question and technique instead of treating multivariable regression and multivariate analysis as interchangeable. Specify the population, period, outcomes, predictors, timing and comparison; address missing data, repeated observations, collinearity, nonlinear relationships, interactions, influential cases and subgroup differences; and report fit, uncertainty and suitable validation. The result depends on the variables and design represented. An adjusted association can support diagnosis or prediction but is not a causal effect unless a separate causal design and assumptions justify that interpretation.
  • Regression testing: Rerun a versioned set of previously passed tests after a model, prompt, data, tool or workflow change to detect lost supported behaviour. Include critical failures and acceptance boundaries and distinguish intended changes from unintended regressions. Passing a regression set protects covered behaviour only; it does not establish that the new feature works or reveal unrepresented failures.

Parent method family

Machine Learning & Predictive Analytics explains how this method connects to adjacent methods and relevant services.

Related service families

These service families contain business questions supported by this method. Service pages link to the wider method family so readers can understand the complete analytical approach.

Explore all Methods & Technologies

Marketways.ai – The Information Highway to your Market!