Zufall und Verteilungen
Mit set.seed() werden Zufallszahlen reproduzierbar:
set.seed(42)
sample(1:6, 5, replace = TRUE)
round(rnorm(5, mean = 100, sd = 15), 1)
round(runif(3), 3)
sample(c("Kopf", "Zahl"), 6, replace = TRUE)
rbinom(5, size = 10, prob = 0.5)
sample(10)
dnorm(0) # Dichte
pnorm(1.96) # P(Z <= 1.96)
qnorm(0.975) # Quantil
dbinom(3, 10, 0.5) # P(X = 3)
pbinom(3, 10, 0.5)
ppois(2, 3)
round(punif(0.25), 2)
choose(5, 2); factorial(5)Ausgabe
[1] 1 5 1 1 2 [1] 100.7 83.4 108.1 108.7 90.1 [1] 0.940 0.978 0.117 [1] "Zahl" "Zahl" "Kopf" "Kopf" "Kopf" "Kopf" [1] 3 5 5 7 5 [1] 5 4 2 9 3 8 1 10 6 7 [1] 0.3989423 [1] 0.9750021 [1] 1.959964 [1] 0.1171875 [1] 0.171875 [1] 0.4231901 [1] 0.25 [1] 10 [1] 120
Das Muster: d (Dichte), p (kumulierte Wahrscheinlichkeit), q (Quantil), r (Zufallszahlen) plus Verteilungsname (norm, binom, pois, unif, exp, t, chisq).
Deskriptive Statistik
x <- c(23, 25, 27, 22, 30, 28, 26, 24, 100)
c(mittel = mean(x), median = median(x), sd = sd(x), iqr = IQR(x))
mean(x, trim = 0.1)
fivenum(x)
boxplot.stats(x)$out # Ausreißer
cor(mtcars$mpg, mtcars$wt)
round(cor(mtcars[, c("mpg", "wt", "hp")]), 2)
var(x) == sd(x)^2
table(cut(x, breaks = c(0, 25, 30, 200)))
round(sapply(mtcars[, c("mpg", "hp")], function(v) c(mean = mean(v), sd = sd(v))), 1)
tapply(mtcars$mpg, mtcars$am, mean)Ausgabe
mittel median sd iqr
33.88889 26.00000 24.91708 4.00000
[1] 33.88889
[1] 22 24 26 28 100
[1] 100
[1] -0.8676594
mpg wt hp
mpg 1.00 -0.87 -0.78
wt -0.87 1.00 0.66
hp -0.78 0.66 1.00
[1] FALSE
(0,25] (25,30] (30,200]
4 4 1
mpg hp
mean 20.1 146.7
sd 6.0 68.6
0 1
17.14737 24.39231Hypothesentests
Ein Test prüft, ob ein beobachteter Unterschied zufällig sein kann. Der p-Wert ist die Wahrscheinlichkeit, bei zutreffender Nullhypothese ein mindestens so extremes Ergebnis zu sehen (üblich: Signifikanz bei p < 0,05).
manuell <- mtcars$mpg[mtcars$am == 0]
automatik <- mtcars$mpg[mtcars$am == 1]
t <- t.test(automatik, manuell) # Welch-Test: Mittelwerte vergleichen
round(t$statistic, 3)
round(t$p.value, 4)
round(t$estimate, 2)
round(t$conf.int, 2)
t.test(mtcars$mpg, mu = 20)$p.value < 0.05 # Einstichproben-Test
tab <- table(mtcars$am, mtcars$cyl)
tab
chisq.test(tab)$p.value < 0.05
round(cor.test(mtcars$mpg, mtcars$wt)$estimate, 3)
shapiro.test(mtcars$mpg)$p.value > 0.05 # Normalverteilung plausibel?
wilcox.test(mpg ~ am, data = mtcars, exact = FALSE)$p.value < 0.05Ausgabe
t
3.767
[1] 0.0014
mean of x mean of y
24.39 17.15
[1] 3.21 11.28
attr(,"conf.level")
[1] 0.95
[1] FALSE
4 6 8
0 3 4 12
1 8 3 2
[1] TRUE
cor
-0.868
[1] TRUE
[1] TRUELineare Regression
modell <- lm(mpg ~ wt, data = mtcars) # mpg erklärt durch Gewicht
round(coef(modell), 3)
round(summary(modell)$r.squared, 3)
round(predict(modell, data.frame(wt = c(2.5, 3.5))), 2)
round(summary(modell)$coefficients, 4)
mehr <- lm(mpg ~ wt + hp, data = mtcars)
round(coef(mehr), 4)
round(AIC(modell), 1) > round(AIC(mehr), 1)
round(head(residuals(modell), 3), 2)
anova(modell)$"Pr(>F)"[1] < 0.001Ausgabe
(Intercept) wt
37.285 -5.344
[1] 0.753
1 2
23.92 18.58
Estimate Std. Error t value Pr(>|t|)
(Intercept) 37.2851 1.8776 19.8576 0
wt -5.3445 0.5591 -9.5590 0
(Intercept) wt hp
37.2273 -3.8778 -0.0318
[1] TRUE
Mazda RX4 Mazda RX4 Wag Datsun 710
-2.28 -0.92 -2.09
[1] TRUEDie Formelsyntax y ~ x1 + x2 wird in vielen Funktionen verwendet (lm, glm, aggregate, boxplot, t.test). Für Klassifikation: glm(..., family = binomial).
logit <- glm(am ~ wt, data = mtcars, family = binomial)
round(coef(logit), 3)
round(predict(logit, data.frame(wt = 2.5), type = "response"), 3)Ausgabe
(Intercept) wt
12.040 -4.024
1
0.879Merke
set.seedfür reproduzierbaren Zufall;d/p/q/r+ Verteilung (norm,binom, ...)- Deskriptiv:
mean,median,sd,quantile,cor,table - Tests:
t.test,chisq.test,cor.test,shapiro.test,wilcox.test(p-Wert < 0,05: signifikant) - Modelle:
lm(y ~ x),summary,predict,glmfür logistische Regression - Statistische Signifikanz ist nicht dasselbe wie praktische Relevanz
Aufgabe
Untersuche mit lm, wie das Gewicht (wt) und die PS (hp) den Verbrauch in mtcars erklären, und vergleiche R².