EMZETT.
Login

Kurz: Mit set.seed() werden Zufallszahlen reproduzierbar:

Teil des Kurses R

Kapitel 7 von 9 im Kurs R (Abschnitt „Statistik“). Mit Fortschritt, Quiz und Zertifikat auf der Lernseite.

Zufall und Verteilungen

Mit set.seed() werden Zufallszahlen reproduzierbar:

set.seed(42)
sample(1:6, 5, replace = TRUE)
round(rnorm(5, mean = 100, sd = 15), 1)
round(runif(3), 3)
sample(c("Kopf", "Zahl"), 6, replace = TRUE)
rbinom(5, size = 10, prob = 0.5)
sample(10)
 
dnorm(0)                       # Dichte
pnorm(1.96)                    # P(Z <= 1.96)
qnorm(0.975)                   # Quantil
dbinom(3, 10, 0.5)             # P(X = 3)
pbinom(3, 10, 0.5)
ppois(2, 3)
round(punif(0.25), 2)
choose(5, 2); factorial(5)

Ausgabe:

[1] 1 5 1 1 2
[1] 100.7  83.4 108.1 108.7  90.1
[1] 0.940 0.978 0.117
[1] "Zahl" "Zahl" "Kopf" "Kopf" "Kopf" "Kopf"
[1] 3 5 5 7 5
 [1]  5  4  2  9  3  8  1 10  6  7
[1] 0.3989423
[1] 0.9750021
[1] 1.959964
[1] 0.1171875
[1] 0.171875
[1] 0.4231901
[1] 0.25
[1] 10
[1] 120

Das Muster: d (Dichte), p (kumulierte Wahrscheinlichkeit), q (Quantil), r (Zufallszahlen) plus Verteilungsname (norm, binom, pois, unif, exp, t, chisq).

Deskriptive Statistik

x <- c(23, 25, 27, 22, 30, 28, 26, 24, 100)
c(mittel = mean(x), median = median(x), sd = sd(x), iqr = IQR(x))
mean(x, trim = 0.1)
fivenum(x)
boxplot.stats(x)$out           # Ausreißer
cor(mtcars$mpg, mtcars$wt)
round(cor(mtcars[, c("mpg", "wt", "hp")]), 2)
var(x) == sd(x)^2
table(cut(x, breaks = c(0, 25, 30, 200)))
round(sapply(mtcars[, c("mpg", "hp")], function(v) c(mean = mean(v), sd = sd(v))), 1)
tapply(mtcars$mpg, mtcars$am, mean)

Ausgabe:

mittel   median       sd      iqr
33.88889 26.00000 24.91708  4.00000
[1] 33.88889
[1]  22  24  26  28 100
[1] 100
[1] -0.8676594
      mpg    wt    hp
mpg  1.00 -0.87 -0.78
wt  -0.87  1.00  0.66
hp  -0.78  0.66  1.00
[1] FALSE
 
  (0,25]  (25,30] (30,200]
       4        4        1
      mpg    hp
mean 20.1 146.7
sd    6.0  68.6
       0        1
17.14737 24.39231

Hypothesentests

Ein Test prüft, ob ein beobachteter Unterschied zufällig sein kann. Der p-Wert ist die Wahrscheinlichkeit, bei zutreffender Nullhypothese ein mindestens so extremes Ergebnis zu sehen (üblich: Signifikanz bei p < 0,05).

manuell <- mtcars$mpg[mtcars$am == 0]
automatik <- mtcars$mpg[mtcars$am == 1]
t <- t.test(automatik, manuell)         # Welch-Test: Mittelwerte vergleichen
round(t$statistic, 3)
round(t$p.value, 4)
round(t$estimate, 2)
round(t$conf.int, 2)
 
t.test(mtcars$mpg, mu = 20)$p.value < 0.05          # Einstichproben-Test
 
tab <- table(mtcars$am, mtcars$cyl)
tab
chisq.test(tab)$p.value < 0.05
round(cor.test(mtcars$mpg, mtcars$wt)$estimate, 3)
shapiro.test(mtcars$mpg)$p.value > 0.05             # Normalverteilung plausibel?
wilcox.test(mpg ~ am, data = mtcars, exact = FALSE)$p.value < 0.05

Ausgabe:

t
3.767
[1] 0.0014
mean of x mean of y
    24.39     17.15
[1]  3.21 11.28
attr(,"conf.level")
[1] 0.95
[1] FALSE
 
     4  6  8
  0  3  4 12
  1  8  3  2
[1] TRUE
   cor
-0.868
[1] TRUE
[1] TRUE

Lineare Regression

modell <- lm(mpg ~ wt, data = mtcars)    # mpg erklärt durch Gewicht
round(coef(modell), 3)
round(summary(modell)$r.squared, 3)
round(predict(modell, data.frame(wt = c(2.5, 3.5))), 2)
round(summary(modell)$coefficients, 4)
mehr <- lm(mpg ~ wt + hp, data = mtcars)
round(coef(mehr), 4)
round(AIC(modell), 1) > round(AIC(mehr), 1)
round(head(residuals(modell), 3), 2)
anova(modell)$"Pr(>F)"[1] < 0.001

Ausgabe:

(Intercept)          wt
     37.285      -5.344
[1] 0.753
    1     2
23.92 18.58
            Estimate Std. Error t value Pr(>|t|)
(Intercept)  37.2851     1.8776 19.8576        0
wt           -5.3445     0.5591 -9.5590        0
(Intercept)          wt          hp
    37.2273     -3.8778     -0.0318
[1] TRUE
    Mazda RX4 Mazda RX4 Wag    Datsun 710
        -2.28         -0.92         -2.09
[1] TRUE

Die Formelsyntax y ~ x1 + x2 wird in vielen Funktionen verwendet (lm, glm, aggregate, boxplot, t.test). Für Klassifikation: glm(..., family = binomial).

logit <- glm(am ~ wt, data = mtcars, family = binomial)
round(coef(logit), 3)
round(predict(logit, data.frame(wt = 2.5), type = "response"), 3)

Ausgabe:

(Intercept)          wt
     12.040      -4.024
    1
0.879

Merke

  • set.seed für reproduzierbaren Zufall; d/p/q/r + Verteilung (norm, binom, …)
  • Deskriptiv: mean, median, sd, quantile, cor, table
  • Tests: t.test, chisq.test, cor.test, shapiro.test, wilcox.test (p-Wert < 0,05: signifikant)
  • Modelle: lm(y ~ x), summary, predict, glm für logistische Regression
  • Statistische Signifikanz ist nicht dasselbe wie praktische Relevanz

Übungsaufgabe

Untersuche mit lm, wie das Gewicht (wt) und die PS (hp) den Verbrauch in mtcars erklären, und vergleiche R².

Quiz zur Selbstkontrolle

Weiter im Kurs

Zurück: Text, Datum und Dateien

Weiter: Grafik, Pakete und Tidyverse

Alle Kapitel: R im Überblick