Skip to content

statsjunk.correlation

pearson

CorrelationResult

Bases: BaseModel

Pearson correlation coefficient, its p-value, and confidence interval.

compute_pearson_correlation

compute_pearson_correlation(
    x: Sequence[float], y: Sequence[float], ci: float = 0.95
) -> CorrelationResult

Compute Pearson's correlation coefficient and its confidence interval.

What this solves

You have two lists of numbers measured on the same subjects (e.g. hours studied and exam score for a group of students) and want to know how strongly they move together — and whether that relationship is likely real or could just be noise. This returns a single number from -1 (perfectly opposite) to +1 (perfectly together), a p-value for whether that number is distinguishable from "no relationship", and a range of plausible values for the true correlation. You don't need to know how the p-value or interval are derived — just that a small p-value and a narrow interval far from zero both mean the relationship is trustworthy.

Parameters:

Name Type Description Default
x Sequence[float]

First variable.

required
y Sequence[float]

Second variable.

required
ci float

Confidence level for the correlation interval. Must be in (0, 1).

0.95

Returns:

Type Description
CorrelationResult

Pearson correlation coefficient (r), p-value, and confidence interval bounds.

Raises:

Type Description
ValueError

If: - x and y have different lengths - fewer than 4 observations are provided - ci is not in (0, 1) - either input is constant

References
  • Pearson, K. (1895). "Note on regression and inheritance in the case of two parents." Proceedings of the Royal Society of London, 58, 240-242.
  • Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.

compute_pearson_from_summary

compute_pearson_from_summary(
    r: float, n: int, ci: float = 0.95
) -> CorrelationResult

Derive a Pearson correlation's p-value and confidence interval from summary statistics (r, n) alone, without the underlying arrays.

What this solves

Sometimes you already know a correlation coefficient and how many observations it was computed from — say, from a report, a previous study, or a database aggregate — but don't have (or don't want to re-fetch) the original raw data. Recomputing the p-value and confidence interval from r and n alone gives the same answer compute_pearson_correlation would have given on the original data, without needing that data at all — for example, a SQL corr() aggregate grouped by category, where you get r and n per group but not the underlying rows.

Parameters:

Name Type Description Default
r float

Pearson correlation coefficient, in [-1, 1].

required
n int

Number of observations the correlation was computed over.

required
ci float

Confidence level for the correlation interval. Must be in (0, 1).

0.95

Returns:

Type Description
CorrelationResult

Same shape as compute_pearson_correlation.

Raises:

Type Description
ValueError

If n < 4, r is outside [-1, 1], or ci is not in (0, 1).

References
  • Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.

spearman

SpearmanCorrelationResult

Bases: BaseModel

Spearman rank correlation coefficient, its p-value, and confidence interval.

compute_spearman_correlation

compute_spearman_correlation(
    x: Sequence[float], y: Sequence[float], ci: float = 0.95
) -> SpearmanCorrelationResult

Compute Spearman's rank correlation coefficient and its confidence interval.

What this solves

Like Pearson correlation, this measures how strongly two variables move together — but based on their ranks rather than their raw values, so it still works when the relationship is consistently increasing or decreasing without being a straight line (e.g. "more of X always means more of Y, but not at a constant rate"), and it isn't thrown off by a few extreme outliers the way Pearson's r can be. Use this instead of compute_pearson_correlation when you care about "do they move in the same direction" more than "is the relationship linear."

Parameters:

Name Type Description Default
x Sequence[float]

First variable.

required
y Sequence[float]

Second variable.

required
ci float

Confidence level for the correlation interval. Must be in (0, 1).

0.95

Returns:

Type Description
SpearmanCorrelationResult

Spearman's rho, p-value, and confidence interval bounds.

Raises:

Type Description
ValueError

If: - x and y have different lengths - fewer than 4 observations are provided - ci is not in (0, 1) - either input is constant

References
  • Spearman, C. (1904). "The proof and measurement of association between two things." American Journal of Psychology, 15(1), 72-101.
  • Fieller, E.C., Hartley, H.O., & Pearson, E.S. (1957). "Tests for rank correlation coefficients. I." Biometrika, 44(3/4), 470-481. (the 1.06 correction to the Fisher z standard error used here)

fisher

fisher_z

fisher_z(r: float) -> float

Fisher z-transformation of a correlation coefficient.

What this solves

Correlation coefficients are squeezed into the range -1 to +1, which makes them awkward to do further statistics with directly (e.g. building a confidence interval, or comparing two correlations). This stretches a correlation onto the full number line, where it behaves much closer to a normal, bell-curve-shaped quantity — the building block behind fisher_z_ci and correlation confidence intervals in general.

z = arctanh(r). Maps r from (-1, 1) to the whole real line, where it is approximately normally distributed — the basis for building confidence intervals and comparing correlations across samples.

Parameters:

Name Type Description Default
r float

Correlation coefficient, in (-1, 1).

required

Returns:

Type Description
float

The Fisher z-transformed value.

References
  • Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.

inverse_fisher_z

inverse_fisher_z(z: float) -> float

Inverse Fisher z-transformation, back to a correlation coefficient.

What this solves

Undoes fisher_z: once you've done statistics on the "stretched out" z-scale (for example, computed a confidence interval around a z-value), this converts the result back to an ordinary correlation coefficient that you can interpret directly.

r = tanh(z). Inverse of fisher_z.

Parameters:

Name Type Description Default
z float

A Fisher z-transformed value.

required

Returns:

Type Description
float

The corresponding correlation coefficient, in (-1, 1).

References
  • Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.

fisher_z_ci

fisher_z_ci(
    r: float, n: int, ci: float = 0.95
) -> tuple[float, float]

Confidence interval for a correlation coefficient via the Fisher z-transform.

What this solves

Gives a range of plausible values for a "true" correlation, given one correlation coefficient computed from a sample of a given size. A wider interval means less certainty (typically from a smaller sample); a narrower interval, further from zero, means the relationship is more likely to be real and not just sampling noise.

Parameters:

Name Type Description Default
r float

Correlation coefficient, in (-1, 1). Callers should handle abs(r) >= 1 themselves — the transform is undefined there.

required
n int

Number of observations the correlation was computed over. Must be greater than 3.

required
ci float

Confidence level. Must be in (0, 1).

0.95

Returns:

Type Description
tuple[float, float]

(low, high) bounds of the confidence interval, in (-1, 1).

Raises:

Type Description
ValueError

If n <= 3 or ci is not in (0, 1).

References
  • Fisher, R.A. (1915). "Frequency distribution of the values of the correlation coefficient in samples of an indefinitely large population." Biometrika, 10(4), 507-521.