Kurz: Mit set.seed() werden Zufallszahlen reproduzierbar:
Teil des Kurses R
Zufall und Verteilungen
Mit set.seed() werden Zufallszahlen reproduzierbar:
set.seed(42)
sample(1:6, 5, replace = TRUE)
round(rnorm(5, mean = 100, sd = 15), 1)
round(runif(3), 3)
sample(c("Kopf", "Zahl"), 6, replace = TRUE)
rbinom(5, size = 10, prob = 0.5)
sample(10)
dnorm(0) # Dichte
pnorm(1.96) # P(Z <= 1.96)
qnorm(0.975) # Quantil
dbinom(3, 10, 0.5) # P(X = 3)
pbinom(3, 10, 0.5)
ppois(2, 3)
round(punif(0.25), 2)
choose(5, 2); factorial(5)Ausgabe:
[1] 1 5 1 1 2
[1] 100.7 83.4 108.1 108.7 90.1
[1] 0.940 0.978 0.117
[1] "Zahl" "Zahl" "Kopf" "Kopf" "Kopf" "Kopf"
[1] 3 5 5 7 5
[1] 5 4 2 9 3 8 1 10 6 7
[1] 0.3989423
[1] 0.9750021
[1] 1.959964
[1] 0.1171875
[1] 0.171875
[1] 0.4231901
[1] 0.25
[1] 10
[1] 120Das Muster: d (Dichte), p (kumulierte Wahrscheinlichkeit), q (Quantil), r (Zufallszahlen) plus Verteilungsname (norm, binom, pois, unif, exp, t, chisq).
Deskriptive Statistik
x <- c(23, 25, 27, 22, 30, 28, 26, 24, 100)
c(mittel = mean(x), median = median(x), sd = sd(x), iqr = IQR(x))
mean(x, trim = 0.1)
fivenum(x)
boxplot.stats(x)$out # Ausreißer
cor(mtcars$mpg, mtcars$wt)
round(cor(mtcars[, c("mpg", "wt", "hp")]), 2)
var(x) == sd(x)^2
table(cut(x, breaks = c(0, 25, 30, 200)))
round(sapply(mtcars[, c("mpg", "hp")], function(v) c(mean = mean(v), sd = sd(v))), 1)
tapply(mtcars$mpg, mtcars$am, mean)Ausgabe:
mittel median sd iqr
33.88889 26.00000 24.91708 4.00000
[1] 33.88889
[1] 22 24 26 28 100
[1] 100
[1] -0.8676594
mpg wt hp
mpg 1.00 -0.87 -0.78
wt -0.87 1.00 0.66
hp -0.78 0.66 1.00
[1] FALSE
(0,25] (25,30] (30,200]
4 4 1
mpg hp
mean 20.1 146.7
sd 6.0 68.6
0 1
17.14737 24.39231Hypothesentests
Ein Test prüft, ob ein beobachteter Unterschied zufällig sein kann. Der p-Wert ist die Wahrscheinlichkeit, bei zutreffender Nullhypothese ein mindestens so extremes Ergebnis zu sehen (üblich: Signifikanz bei p < 0,05).
manuell <- mtcars$mpg[mtcars$am == 0]
automatik <- mtcars$mpg[mtcars$am == 1]
t <- t.test(automatik, manuell) # Welch-Test: Mittelwerte vergleichen
round(t$statistic, 3)
round(t$p.value, 4)
round(t$estimate, 2)
round(t$conf.int, 2)
t.test(mtcars$mpg, mu = 20)$p.value < 0.05 # Einstichproben-Test
tab <- table(mtcars$am, mtcars$cyl)
tab
chisq.test(tab)$p.value < 0.05
round(cor.test(mtcars$mpg, mtcars$wt)$estimate, 3)
shapiro.test(mtcars$mpg)$p.value > 0.05 # Normalverteilung plausibel?
wilcox.test(mpg ~ am, data = mtcars, exact = FALSE)$p.value < 0.05Ausgabe:
t
3.767
[1] 0.0014
mean of x mean of y
24.39 17.15
[1] 3.21 11.28
attr(,"conf.level")
[1] 0.95
[1] FALSE
4 6 8
0 3 4 12
1 8 3 2
[1] TRUE
cor
-0.868
[1] TRUE
[1] TRUELineare Regression
modell <- lm(mpg ~ wt, data = mtcars) # mpg erklärt durch Gewicht
round(coef(modell), 3)
round(summary(modell)$r.squared, 3)
round(predict(modell, data.frame(wt = c(2.5, 3.5))), 2)
round(summary(modell)$coefficients, 4)
mehr <- lm(mpg ~ wt + hp, data = mtcars)
round(coef(mehr), 4)
round(AIC(modell), 1) > round(AIC(mehr), 1)
round(head(residuals(modell), 3), 2)
anova(modell)$"Pr(>F)"[1] < 0.001Ausgabe:
(Intercept) wt
37.285 -5.344
[1] 0.753
1 2
23.92 18.58
Estimate Std. Error t value Pr(>|t|)
(Intercept) 37.2851 1.8776 19.8576 0
wt -5.3445 0.5591 -9.5590 0
(Intercept) wt hp
37.2273 -3.8778 -0.0318
[1] TRUE
Mazda RX4 Mazda RX4 Wag Datsun 710
-2.28 -0.92 -2.09
[1] TRUEDie Formelsyntax y ~ x1 + x2 wird in vielen Funktionen verwendet (lm, glm, aggregate, boxplot, t.test). Für Klassifikation: glm(..., family = binomial).
logit <- glm(am ~ wt, data = mtcars, family = binomial)
round(coef(logit), 3)
round(predict(logit, data.frame(wt = 2.5), type = "response"), 3)Ausgabe:
(Intercept) wt
12.040 -4.024
1
0.879Merke
set.seedfür reproduzierbaren Zufall;d/p/q/r+ Verteilung (norm,binom, …)- Deskriptiv:
mean,median,sd,quantile,cor,table - Tests:
t.test,chisq.test,cor.test,shapiro.test,wilcox.test(p-Wert < 0,05: signifikant) - Modelle:
lm(y ~ x),summary,predict,glmfür logistische Regression - Statistische Signifikanz ist nicht dasselbe wie praktische Relevanz
Übungsaufgabe
Untersuche mit lm, wie das Gewicht (wt) und die PS (hp) den Verbrauch in mtcars erklären, und vergleiche R².
Quiz zur Selbstkontrolle
Was bedeutet ein p-Wert unter 0,05 üblicherweise?
- Das Ergebnis gilt als statistisch signifikant (richtig)
- Die Hypothese ist bewiesen
- Die Daten sind falsch
- Der Effekt ist groß
Wofür steht die Formel y ~ x?
- y wird durch x erklärt (richtig)
- y ungefähr x
- y minus x
- y hoch x
Was bewirkt set.seed(42)?
- Zufallszahlen werden reproduzierbar (richtig)
- Zufallszahlen werden schneller
- Der Zufall wird abgeschaltet
- Das Programm läuft nur einmal
Weiter im Kurs
Zurück: Text, Datum und Dateien
Weiter: Grafik, Pakete und Tidyverse
Alle Kapitel: R im Überblick