Logistic regression is one of the most used statistical techniques for explaining the probabilistic behavior of a given phenomenon. Data separation is a frequent problem in this model, as successes appear separated from failures and make it impossible to find the maximum likelihood estimators. Objective: to present a revision and a solution to the problem, and to compare it with other solutions. Methodology: a simulation of the logistic model and an estimation of the parameters’ bias using the proposed classical and Bayesian solution with fictitious observations, as well as the Firth method. Results: the bias found is lower when the pair of fictitious observations are generated using the Bayesian method. An example about the age at which menarche occurs is presented. Discussion: an appropriate solution to the problem of separation is provided using a simulation in a simple logistic model. Conclusions: the generation of fictitious observations within the separation region is recommended, and the best solution method is based on Bayesian theory, which achieves convergence of the parameters of the logistic model.
|Translated title of the contribution
|The problem of separation in logistic regression, a solution and an application
|Number of pages
|Revista Facultad Nacional de Salud Publica
|Published - Sept 2011
Bibliographical notePublisher Copyright:
© Universidad de Antioquia. All Rights Reserved.
All Science Journal Classification (ASJC) codes
- Health Policy
- Public Health, Environmental and Occupational Health
- Health Information Management