statsjunk.regression¶
simple ¶
RegressionResult ¶
Bases: BaseModel
Result of a simple linear regression, with slope diagnostics.
compute_linear_regression ¶
compute_linear_regression(
x: Sequence[float], y: Sequence[float], ci: float = 0.95
) -> RegressionResult
Fit a simple linear regression model and its slope diagnostics.
What this solves
You have one variable you think predicts or explains another (e.g. years
of experience predicting salary), and want the straight line that best
fits the data, plus a sense of how reliable that line is. This fits the
best-fitting line, tells you how much of the variation in y it
explains (r_squared), and gives a confidence interval and p-value for
the slope — so you can judge whether the relationship is likely real
(slope clearly different from zero) or could just be noise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
Sequence[float]
|
Independent variable. |
required |
y
|
Sequence[float]
|
Dependent variable. |
required |
ci
|
float
|
Confidence level for the slope interval. Must be in (0, 1). |
0.95
|
Returns:
| Type | Description |
|---|---|
RegressionResult
|
Fitted |
Raises:
| Type | Description |
|---|---|
ValueError
|
If: - x and y have different lengths - fewer than 3 observations are provided - ci is not in (0, 1) - either input is constant |
References
- Draper, N.R. & Smith, H. (1998). Applied Regression Analysis (3rd ed.). Wiley.
multiple ¶
MultipleRegressionResult ¶
Bases: BaseModel
Result of an ordinary-least-squares multiple linear regression.
Coefficient ¶
Bases: BaseModel
One term of a fitted multiple linear regression.
coef is on the regressor's own scale; coef_std is the standardized
(beta) coefficient — the effect in standard deviations of y per standard
deviation of the regressor — so magnitudes are comparable across regressors
on different scales. Both are None for the intercept, as is vif.
compute_multiple_regression ¶
compute_multiple_regression(
x: Sequence[Sequence[float]],
y: Sequence[float],
names: Sequence[str] | None = None,
ci: float = 0.95,
) -> MultipleRegressionResult
Fit a multiple linear regression by ordinary least squares.
What this solves
You have several variables you think together predict or explain an
outcome (e.g. education, experience, and age predicting salary), and
want to know each one's individual effect while accounting for the
others. This fits the best-fitting linear combination, reports each
predictor's effect size, standard error, and significance, and flags
predictors that are too similar to each other to separate reliably
(vif, variance inflation factor — a large value there is a warning
sign that two or more of your predictors carry redundant information).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
Sequence[Sequence[float]]
|
Design matrix of shape |
required |
y
|
Sequence[float]
|
Dependent variable, one value per observation. |
required |
names
|
Sequence[str]
|
Label for each regressor column; defaults to |
None
|
ci
|
float
|
Confidence level for the per-coefficient intervals. Must be in (0, 1). |
0.95
|
Returns:
| Type | Description |
|---|---|
MultipleRegressionResult
|
Per-coefficient estimates (raw and standardized), standard errors,
two-sided Wald p-values, |
Raises:
| Type | Description |
|---|---|
ValueError
|
If:
- x is not 2-D, or x and y disagree on the number of observations
- there are no regressors, or fewer than |
References
- Draper, N.R. & Smith, H. (1998). Applied Regression Analysis (3rd ed.). Wiley.
- Marquardt, D.W. (1970). "Generalized Inverses, Ridge Regression, Biased Linear Estimation, and Nonlinear Estimation." Technometrics, 12(3), 591-612. (variance inflation factor)