Categorical Data Analysis by Example
Autor Graham J G Uptonen Limba Engleză Hardback – 14 noi 2016
Preț: 544.83 lei
Preț vechi: 844.41 lei
-35%
Puncte Express: 817
Carte indisponibilă temporar
Doresc să fiu notificat când acest titlu va fi disponibil:
Se trimite...
Specificații
ISBN-13: 9781119307860
ISBN-10: 1119307864
Pagini: 224
Dimensiuni: 155 x 236 x 18 mm
Greutate: 0.48 kg
Editura: Wiley
Locul publicării:Hoboken, United States
ISBN-10: 1119307864
Pagini: 224
Dimensiuni: 155 x 236 x 18 mm
Greutate: 0.48 kg
Editura: Wiley
Locul publicării:Hoboken, United States
Public țintă
Primary: Students in statistics and researchers in other disciplines, especially the social sciences, who use categorical data.Secondary: Practitioners in market research, medicine and other fields.
Notă biografică
GRAHAM J. G. UPTON is formerly Professor of Applied Statistics, Department of Mathematical Sciences, University of Essex. Dr. Upton is author of The Analysis of Cross-tabulated Data (1978) and joint author of Spatial Data Analysis by Example (2 volumes, 1995), both published by Wiley. He is the lead author of The Oxford Dictionary of Statistics (OUP, 2014). His books have been translated into Japanese, Russian, and Welsh.
Cuprins
Preface xi
Acknowledgments xiii
1 Introduction 1
1.1 What are categorical data? 1
1.2 A typical data set 2
1.3 Visualisation and crosstabulation 3
1.4 Samples, populations, and random variation 4
1.5 Proportion, probability and conditional probability 5
1.6 Probability distributions 6
2 Estimation and inference for categorical data 11
2.1 Goodness of fit 11
2.2 Hypothesis tests for a binomial proportion (large sample) 14
2.3 Hypothesis tests for a binomial proportion (small sample) 16
2.4 Interval estimates for a binomial proportion 18
3 The 2 X 2 contingency table 23
3.1 Introduction 23
3.2 Fisher's exact test (for independence) 24
3.3 Testing independence with large cell frequencies 27
3.4 The 2 X 2 table in a medical context 29
3.5 Measuring lack of independence (comparing proportions) 31
4 The I x J contingency table 37
4.1 Notation 37
4.2 Independence in the I X J contingency table 38
4.3 Partitioning 42
4.4 Graphical displays 44
4.5 Testing independence with ordinal variables 46
5 The exponential family 51
5.1 Introduction 51
5.2 The exponential family 52
5.3 Components of a general linear model 53
5.4 Estimation 54
6 A model taxonomy 57
6.1 Underlying questions 57
6.2 Identifying the type of model 58
7 The 2 X J contingency table 61
7.1 A problem with X2 (and G2) 61
7.2 Using the logit 62
7.3 Individual data and grouped data 64
7.4 Precision, confidence intervals, and prediction intervals 69
7.5 Logistic regression with a categorical explanatory variable 70
8 Logistic regression with several explanatory variables 77
8.1 Degrees of freedom when there are no interactions 77
8.2 Getting a feel for the data 79
8.3 Models with two variable interactions 81
9 Model selection and diagnostics 85
9.1 Introduction 85
9.2 Notation for interactions and for models 87
9.3 Stepwise methods for model selection using G2 89
9.4 AIC and related measures 93
9.5 The problem caused by rare combinations of events 95
9.6 Simplicity versus accuracy 98
9.7 DFBETAS 100
10 Multinomial logistic regression 103
10.1 A single continuous explanatory variable 103
10.2 Nominal categorical explanatory variables 106
10.3 Models for an ordinal response variable 108
11 Log-linear models for I X J tables 119
11.1 The saturated model 119
11.2 The independence model for an I X J table 125
12 Log-linear models for I X J X K tables 129
12.1 Mutual independence: A=B=C 131
12.2 The model AB=C 131
12.3 Conditional independence and independence 133
12.4 The model AB=AC 134
12.5 The models AB=AC=BC and ABC 135
12.6 Simpson's paradox 135
12.7 Connection between log-linear models and logistic regression 137
13 Implications and uses of Birch's result 141
13.1 Birch's result 141
13.2 Iterative scaling 142
13.3 The hierarchy constraint 143
13.4 Inclusion of the all-factor interaction 144
13.5 Mostellerising 145
14 Model selection for log-linear models 149
14.1 Three variables 150
14.2 More than three variables 153
15 Incomplete tables, dummy variables, and outliers 157
15.1 Incomplete tables 157
15.2 Quasi-independence 159
15.3 Dummy variables 159
15.4 Detection of outliers 160
16 Panel data and repeated measures 165
16.1 The mover-stayer model 166
16.2 The loyalty model 168
16.3 Symmetry 169
16.4 Quasi-symmetry 170
16.5 The loyalty-distance model 172
A R code for Cobweb function 175
Index 179
Author Index 183
Index of Examples 185
Acknowledgments xiii
1 Introduction 1
1.1 What are categorical data? 1
1.2 A typical data set 2
1.3 Visualisation and crosstabulation 3
1.4 Samples, populations, and random variation 4
1.5 Proportion, probability and conditional probability 5
1.6 Probability distributions 6
2 Estimation and inference for categorical data 11
2.1 Goodness of fit 11
2.2 Hypothesis tests for a binomial proportion (large sample) 14
2.3 Hypothesis tests for a binomial proportion (small sample) 16
2.4 Interval estimates for a binomial proportion 18
3 The 2 X 2 contingency table 23
3.1 Introduction 23
3.2 Fisher's exact test (for independence) 24
3.3 Testing independence with large cell frequencies 27
3.4 The 2 X 2 table in a medical context 29
3.5 Measuring lack of independence (comparing proportions) 31
4 The I x J contingency table 37
4.1 Notation 37
4.2 Independence in the I X J contingency table 38
4.3 Partitioning 42
4.4 Graphical displays 44
4.5 Testing independence with ordinal variables 46
5 The exponential family 51
5.1 Introduction 51
5.2 The exponential family 52
5.3 Components of a general linear model 53
5.4 Estimation 54
6 A model taxonomy 57
6.1 Underlying questions 57
6.2 Identifying the type of model 58
7 The 2 X J contingency table 61
7.1 A problem with X2 (and G2) 61
7.2 Using the logit 62
7.3 Individual data and grouped data 64
7.4 Precision, confidence intervals, and prediction intervals 69
7.5 Logistic regression with a categorical explanatory variable 70
8 Logistic regression with several explanatory variables 77
8.1 Degrees of freedom when there are no interactions 77
8.2 Getting a feel for the data 79
8.3 Models with two variable interactions 81
9 Model selection and diagnostics 85
9.1 Introduction 85
9.2 Notation for interactions and for models 87
9.3 Stepwise methods for model selection using G2 89
9.4 AIC and related measures 93
9.5 The problem caused by rare combinations of events 95
9.6 Simplicity versus accuracy 98
9.7 DFBETAS 100
10 Multinomial logistic regression 103
10.1 A single continuous explanatory variable 103
10.2 Nominal categorical explanatory variables 106
10.3 Models for an ordinal response variable 108
11 Log-linear models for I X J tables 119
11.1 The saturated model 119
11.2 The independence model for an I X J table 125
12 Log-linear models for I X J X K tables 129
12.1 Mutual independence: A=B=C 131
12.2 The model AB=C 131
12.3 Conditional independence and independence 133
12.4 The model AB=AC 134
12.5 The models AB=AC=BC and ABC 135
12.6 Simpson's paradox 135
12.7 Connection between log-linear models and logistic regression 137
13 Implications and uses of Birch's result 141
13.1 Birch's result 141
13.2 Iterative scaling 142
13.3 The hierarchy constraint 143
13.4 Inclusion of the all-factor interaction 144
13.5 Mostellerising 145
14 Model selection for log-linear models 149
14.1 Three variables 150
14.2 More than three variables 153
15 Incomplete tables, dummy variables, and outliers 157
15.1 Incomplete tables 157
15.2 Quasi-independence 159
15.3 Dummy variables 159
15.4 Detection of outliers 160
16 Panel data and repeated measures 165
16.1 The mover-stayer model 166
16.2 The loyalty model 168
16.3 Symmetry 169
16.4 Quasi-symmetry 170
16.5 The loyalty-distance model 172
A R code for Cobweb function 175
Index 179
Author Index 183
Index of Examples 185