statsjunk.correlation¶
pearson ¶
CorrelationResult ¶
Bases: BaseModel
Pearson correlation coefficient, its p-value, and confidence interval.
compute_pearson_correlation ¶
compute_pearson_correlation(
x: Sequence[float], y: Sequence[float], ci: float = 0.95
) -> CorrelationResult
Compute Pearson's correlation coefficient and its confidence interval.
What this solves
You have two lists of numbers measured on the same subjects (e.g. hours studied and exam score for a group of students) and want to know how strongly they move together — and whether that relationship is likely real or could just be noise. This returns a single number from -1 (perfectly opposite) to +1 (perfectly together), a p-value for whether that number is distinguishable from "no relationship", and a range of plausible values for the true correlation. You don't need to know how the p-value or interval are derived — just that a small p-value and a narrow interval far from zero both mean the relationship is trustworthy.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
Sequence[float]
|
First variable. |
required |
y
|
Sequence[float]
|
Second variable. |
required |
ci
|
float
|
Confidence level for the correlation interval. Must be in (0, 1). |
0.95
|
Returns:
| Type | Description |
|---|---|
CorrelationResult
|
Pearson correlation coefficient ( |
Raises:
| Type | Description |
|---|---|
ValueError
|
If: - x and y have different lengths - fewer than 4 observations are provided - ci is not in (0, 1) - either input is constant |
References
- Pearson, K. (1895). "Note on regression and inheritance in the case of two parents." Proceedings of the Royal Society of London, 58, 240-242.
- Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.
compute_pearson_from_summary ¶
compute_pearson_from_summary(
r: float, n: int, ci: float = 0.95
) -> CorrelationResult
Derive a Pearson correlation's p-value and confidence interval from summary statistics (r, n) alone, without the underlying arrays.
What this solves
Sometimes you already know a correlation coefficient and how many
observations it was computed from — say, from a report, a previous
study, or a database aggregate — but don't have (or don't want to
re-fetch) the original raw data. Recomputing the p-value and confidence
interval from r and n alone gives the same answer compute_pearson_correlation
would have given on the original data, without needing that data at all —
for example, a SQL corr() aggregate grouped by category, where you get
r and n per group but not the underlying rows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
r
|
float
|
Pearson correlation coefficient, in [-1, 1]. |
required |
n
|
int
|
Number of observations the correlation was computed over. |
required |
ci
|
float
|
Confidence level for the correlation interval. Must be in (0, 1). |
0.95
|
Returns:
| Type | Description |
|---|---|
CorrelationResult
|
Same shape as |
Raises:
| Type | Description |
|---|---|
ValueError
|
If n < 4, r is outside [-1, 1], or ci is not in (0, 1). |
References
- Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.
spearman ¶
SpearmanCorrelationResult ¶
Bases: BaseModel
Spearman rank correlation coefficient, its p-value, and confidence interval.
compute_spearman_correlation ¶
compute_spearman_correlation(
x: Sequence[float], y: Sequence[float], ci: float = 0.95
) -> SpearmanCorrelationResult
Compute Spearman's rank correlation coefficient and its confidence interval.
What this solves
Like Pearson correlation, this measures how strongly two variables move
together — but based on their ranks rather than their raw values, so
it still works when the relationship is consistently increasing or
decreasing without being a straight line (e.g. "more of X always means
more of Y, but not at a constant rate"), and it isn't thrown off by a
few extreme outliers the way Pearson's r can be. Use this instead of
compute_pearson_correlation when you care about "do they move in the
same direction" more than "is the relationship linear."
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
x
|
Sequence[float]
|
First variable. |
required |
y
|
Sequence[float]
|
Second variable. |
required |
ci
|
float
|
Confidence level for the correlation interval. Must be in (0, 1). |
0.95
|
Returns:
| Type | Description |
|---|---|
SpearmanCorrelationResult
|
Spearman's rho, p-value, and confidence interval bounds. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If: - x and y have different lengths - fewer than 4 observations are provided - ci is not in (0, 1) - either input is constant |
References
- Spearman, C. (1904). "The proof and measurement of association between two things." American Journal of Psychology, 15(1), 72-101.
- Fieller, E.C., Hartley, H.O., & Pearson, E.S. (1957). "Tests for rank correlation coefficients. I." Biometrika, 44(3/4), 470-481. (the 1.06 correction to the Fisher z standard error used here)
fisher ¶
fisher_z ¶
fisher_z(r: float) -> float
Fisher z-transformation of a correlation coefficient.
What this solves
Correlation coefficients are squeezed into the range -1 to +1, which
makes them awkward to do further statistics with directly (e.g. building
a confidence interval, or comparing two correlations). This stretches a
correlation onto the full number line, where it behaves much closer to a
normal, bell-curve-shaped quantity — the building block behind
fisher_z_ci and correlation confidence intervals in general.
z = arctanh(r). Maps r from (-1, 1) to the whole real line,
where it is approximately normally distributed — the basis for building
confidence intervals and comparing correlations across samples.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
r
|
float
|
Correlation coefficient, in |
required |
Returns:
| Type | Description |
|---|---|
float
|
The Fisher z-transformed value. |
References
- Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.
inverse_fisher_z ¶
inverse_fisher_z(z: float) -> float
Inverse Fisher z-transformation, back to a correlation coefficient.
What this solves
Undoes fisher_z: once you've done statistics on the "stretched out"
z-scale (for example, computed a confidence interval around a z-value),
this converts the result back to an ordinary correlation coefficient
that you can interpret directly.
r = tanh(z). Inverse of fisher_z.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
z
|
float
|
A Fisher z-transformed value. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The corresponding correlation coefficient, in |
References
- Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.
fisher_z_ci ¶
fisher_z_ci(
r: float, n: int, ci: float = 0.95
) -> tuple[float, float]
Confidence interval for a correlation coefficient via the Fisher z-transform.
What this solves
Gives a range of plausible values for a "true" correlation, given one correlation coefficient computed from a sample of a given size. A wider interval means less certainty (typically from a smaller sample); a narrower interval, further from zero, means the relationship is more likely to be real and not just sampling noise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
r
|
float
|
Correlation coefficient, in |
required |
n
|
int
|
Number of observations the correlation was computed over. Must be greater than 3. |
required |
ci
|
float
|
Confidence level. Must be in (0, 1). |
0.95
|
Returns:
| Type | Description |
|---|---|
tuple[float, float]
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If n <= 3 or ci is not in (0, 1). |
References
- Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.