Picture a bank trying to decide whether a loan applicant will default, or a hospital trying to figure out which patients are at risk of a heart attack. In both cases, the outcome isn’t a number on a continuous scale. It’s a category: default or no default, heart attack or no heart attack. When your dependent variable is grouped like this instead of continuous, ordinary regression breaks down, and two multivariate techniques step in to fill the gap: Multiple Discriminant Analysis (MDA) and logistic regression. Both are built to handle categorical outcomes, but they approach the problem from different angles, rest on different assumptions, and answer slightly different questions. Knowing which one fits your data and your research goal can save you from a flawed model and a misleading conclusion.
Table of Contents
- What is multiple discriminant analysis?
- How MDA builds its discriminant function
- What MDA expects from your data
- What is logistic regression?
- The logit function and probability
- Odds ratios: the practical payoff
- Discriminant analysis vs logistic regression: the key differences
- Assumptions about the data
- What each method is actually built to tell you
- Robustness in practice
- When should you pick MDA over logistic regression?
- Practical applications in research
- Medical research: predicting disease risk
- Finance: predicting bankruptcy with the Altman Z-score
- Sociology and market research
- What do you think?
What is multiple discriminant analysis?
Multiple Discriminant Analysis is the technique of choice when your dependent variable is nonmetric, meaning it’s dichotomous (such as pass/fail or male/female) or multi-chotomous (such as high/medium/low income groups). Rather than predicting a numeric value, MDA tries to classify each observation into one of the existing groups based on a set of independent variables.
How MDA builds its discriminant function
MDA works by combining the independent variables into a new composite variable called a discriminant score, or a variate. This variate is constructed specifically to maximise the difference between group means while minimising the variation within each group. In simpler terms, it draws the sharpest possible line between categories using the information available. This idea traces back to Ronald Fisher’s original work on discriminant functions, which forms the mathematical backbone of the technique still used in modern classification models today.
What MDA expects from your data
MDA isn’t assumption-free. It performs best when the independent variables are continuous, multivariate normal, and when the covariance matrices across groups are roughly equal. When these conditions hold, MDA tends to produce highly efficient estimates. But real-world data, especially in business and social science research, rarely lines up this neatly. Independent variables are often a mix of categorical and continuous measures, which is precisely where the assumptions of MDA start to strain.
What is logistic regression?
Logistic regression takes a different route to the same broad problem: explaining and predicting categorical outcomes. It’s most commonly used when the dependent variable is binary, such as whether a customer buys a product (yes/no) or whether a patient develops a disease (yes/no). Unlike MDA, logistic regression doesn’t try to find the sharpest separating line between groups. Instead, it models the probability that an observation falls into a particular category, based on its independent variables.
The logit function and probability
The technique gets its name from the logit function, which converts a linear combination of the independent variables into a probability that always falls between 0 and 1. This is important because plain linear regression can produce probability estimates outside that range, which makes no practical sense. So, instead of asking “which group does this belong to?” the way MDA does, logistic regression asks “how does the probability of belonging to a group change as each independent variable changes?” A classic example is examining how the probability of developing lung cancer shifts with every additional pack of cigarettes smoked per day, holding other factors like body weight constant. This framing is why the method is so heavily used in clinical research, where researchers need to isolate the individual contribution of each risk factor.
Odds ratios: the practical payoff
One of the biggest reasons logistic regression is popular is the odds ratio. Once you run a logistic regression, exponentiating the coefficient for a variable gives you its odds ratio, a single number that tells you how much the odds of the outcome change for a one-unit increase in that variable. An odds ratio of 2 for smoking, for instance, means smokers have roughly twice the odds of the outcome compared to non-smokers, holding other variables constant. This interpretability is a major reason the method is standard in epidemiology, clinical trials, and case-control studies, where researchers need to communicate risk in a way that’s intuitive to both scientists and policymakers.
Discriminant analysis vs logistic regression: the key differences
Both techniques evaluate the relationship between a set of covariates and a categorical outcome, and both are extensively used in medical and sociological research. But the differences matter once you start applying them.
Assumptions about the data
MDA assumes the independent variables are multivariate normal with equal covariance matrices across groups. Logistic regression carries almost no such baggage. It doesn’t require normally distributed predictors, and it can comfortably handle a mix of categorical, ordinal, and continuous independent variables. This is one reason logistic regression has become the more popular default across many fields, particularly whenever predictors don’t behave the way MDA would like them to. A well-cited comparison from a statistical review of the two approaches notes that discriminant analysis estimators are preferable only when the underlying populations genuinely are normal with identical covariance structures. Outside that narrow condition, logistic regression tends to be the safer, more robust choice.
What each method is actually built to tell you
MDA is built to answer a classification question: given a new observation, which group does it most likely belong to, and which dimensions best separate the groups? It’s a technique oriented toward prediction of group membership. Logistic regression, by contrast, is built to answer a probability and explanatory question: as a specific independent variable changes, how does the likelihood of the outcome shift, all else being equal? If your research goal is pure classification accuracy, MDA can be a strong fit when its assumptions hold. If your goal is to explain and quantify the effect of specific risk factors, logistic regression’s odds ratios usually give you a clearer story.
Robustness in practice
Interestingly, empirical comparisons often find the two methods perform similarly in practice, even when textbook assumptions favour one over the other. A study comparing the two approaches for sex estimation from skeletal measurements found that discriminant function analysis and logistic regression produced very similar classification accuracy across multiple conditions, with neither method holding a decisive edge. This suggests that while the theoretical assumptions differ, the practical gap between the two techniques can be narrower than it first appears, especially with well-behaved data and reasonable sample sizes.
When should you pick MDA over logistic regression?
There isn’t a universal rule, but a few practical signals help guide the decision.
Choose MDA when: your independent variables are continuous and reasonably normally distributed, your sample size is comfortably large relative to the number of predictors, and your primary interest is in accurately classifying new observations into existing groups.
Choose logistic regression when: your predictors are a mix of categorical and continuous variables, you can’t verify normality assumptions, or your main interest lies in interpreting how each predictor individually shifts the probability of an outcome through odds ratios.
In situations where you’re unsure, it often makes sense to run both and compare classification accuracy and interpretability. Many statisticians now default to logistic regression precisely because of its flexibility with different types of data, even when discriminant analysis would technically be the “correct” choice under strict textbook conditions.
Practical applications in research
Both techniques show up constantly across medical, financial, and social science research, wherever the outcome of interest is a category rather than a number.
Medical research: predicting disease risk
In clinical studies, researchers frequently want to know how factors like body weight, calorie intake, fat intake, and age jointly influence the odds of an event such as a heart attack (yes vs. no). Logistic regression is especially well suited here because it produces individual odds ratios for each risk factor while adjusting for the others, letting clinicians say, for example, that a unit increase in a particular measure raises the odds of the event by a specific, quantifiable amount.
Finance: predicting bankruptcy with the Altman Z-score
One of the most famous real-world applications of MDA is Edward Altman’s Z-score model, developed in 1968 to predict corporate bankruptcy. Altman combined five financial ratios, including working capital to total assets and earnings before interest and tax to total assets, into a single discriminant score. Firms scoring below a certain threshold were classified as high bankruptcy risk, while those above another threshold were considered financially healthy. This model remains a foundational example of how MDA reduces multiple financial indicators into one classification score, and it continues to influence credit risk assessment used by banks and rating agencies today.
Sociology and market research
Beyond medicine and finance, both techniques are widely used to classify consumers, voters, or survey respondents into meaningful segments. A market researcher might use MDA to understand which purchasing habits and demographic traits best separate loyal customers from occasional buyers, while a sociologist might use logistic regression to study how income, education, and age affect the probability of a particular voting choice. In both cases, the underlying goal is the same: making sense of a categorical outcome using multiple explanatory variables, just approached from different statistical angles.
What do you think?
What do you think? If you were analysing a dataset where you weren’t sure whether your predictors met the normality assumptions required by MDA, would you still attempt discriminant analysis, or default straight to logistic regression for safety? And in your field of interest, do you see more value in a technique that classifies group membership precisely, or one that explains how each factor shifts the probability of an outcome?
References
- https://en.wikipedia.org/wiki/Altman_Z-score
- https://www.annalsthoracicsurgery.org/article/S0003-4975(02)04683-0/fulltext
- https://www.2minutemedicine.com/odds-ratio/
- https://people.stat.sc.edu/hoyen/PastTeaching/STAT705-2019/Presentation/LogitOrLDA.pdf
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7497157/
Leave a Reply