beginning r -...

30

Upload: lebao

Post on 23-Aug-2018

214 views

Category:

Documents


0 download

TRANSCRIPT

Page 1: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how
Page 2: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how
Page 3: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

Beginning R

intRoduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xxi

chapteR 1 Introducing R: What It Is and How to Get It . . . . . . . . . . . . . . . . . . . . . . . . . . 1

chapteR 2 Starting Out: Becoming Familiar with R . . . . . . . . . . . . . . . . . . . . . . . . . . . 25

chapteR 3 Starting Out: Working With Objects . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 65

chapteR 4 Data: Descriptive Statistics and Tabulation . . . . . . . . . . . . . . . . . . . . . . . 107

chapteR 5 Data: Distribution . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 151

chapteR 6 Simple Hypothesis Testing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 181

chapteR 7 Introduction to Graphical Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 215

chapteR 8 Formula Notation and Complex Statistics . . . . . . . . . . . . . . . . . . . . . . . . 263

chapteR 9 Manipulating Data and Extracting Components . . . . . . . . . . . . . . . . . . 295

chapteR 10 Regression (Linear Modeling) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 327

chapteR 11 More About Graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 363

chapteR 12 Writing Your Own Scripts: Beginning to Program . . . . . . . . . . . . . . . . . 415

appendix Answers to Exercises . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 433

index . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .461

Page 4: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how
Page 5: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

Beginning

Rthe StatiStical pRogRamming language

Page 6: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how
Page 7: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

Beginning

Rthe StatiStical pRogRamming language

Mark Gardener

Page 8: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

Beginning R: the Statistical programming language

Published by John Wiley & Sons, Inc. 10475 Crosspoint Boulevard Indianapolis, IN 46256 www.wiley.com

Copyright © 2012 by John Wiley & Sons, Inc., Indianapolis, Indiana

Published simultaneously in Canada

ISBN: 978-1-118-16430-3

ISBN: 978-1-118-22616-2 (ebk)

ISBN: 978-1-118-23937-7 (ebk)

ISBN: 978-1-118-26412-6 (ebk)

Manufactured in the United States of America

10 9 8 7 6 5 4 3 2 1

No part of this publication may be reproduced, stored in a retrieval system or transmitted in any form or by any means, electronic, mechanical, photocopying, recording, scanning or otherwise, except as permitted under Sections 107 or 108 of the 1976 United States Copyright Act, without either the prior written permission of the Publisher, or authorization through payment of the appropriate per-copy fee to the Copyright Clearance Center, 222 Rosewood Drive, Danvers, MA 01923, (978) 750-8400, fax (978) 646-8600. Requests to the Publisher for permission should be addressed to the Permissions Department, John Wiley & Sons, Inc., 111 River Street, Hoboken, NJ 07030, (201) 748-6011, fax (201) 748-6008, or online at http://www.wiley.com/go/permissions.

Limit of Liability/Disclaimer of Warranty: The publisher and the author make no representations or warranties with respect to the accuracy or completeness of the contents of this work and specifically disclaim all warranties, including without limitation warranties of fitness for a particular purpose. No warranty may be created or extended by sales or promotional materials. The advice and strategies contained herein may not be suitable for every situation. This work is sold with the understanding that the publisher is not engaged in rendering legal, accounting, or other professional services. If professional assistance is required, the services of a competent professional person should be sought. Neither the publisher nor the author shall be liable for damages arising herefrom. The fact that an organization or Web site is referred to in this work as a citation and/or a potential source of further information does not mean that the author or the publisher endorses the information the organization or Web site may provide or recommendations it may make. Further, readers should be aware that Internet Web sites listed in this work may have changed or disappeared between when this work was written and when it is read.

For general information on our other products and services please contact our Customer Care Department within the United States at (877) 762-2974, outside the United States at (317) 572-3993 or fax (317) 572-4002.

Wiley publishes in a variety of print and electronic formats and by print-on-demand. Some material included with standard print versions of this book may not be included in e-books or in print-on-demand. If this book refers to media such as a CD or DVD that is not included in the version you purchased, you may download this material at http://booksupport .wiley.com. For more information about Wiley products, visit www.wiley.com.

Library of Congress Control Number: 2012937909

Trademarks: Wiley, the Wiley logo, Wrox, the Wrox logo, Programmer to Programmer, and related trade dress are trademarks or registered trademarks of John Wiley & Sons, Inc. and/or its affiliates, in the United States and other countries, and may not be used without written permission. All other trademarks are the property of their respective owners. John Wiley & Sons, Inc., is not associated with any product or vendor mentioned in this book.

Page 9: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

It is much easier to be critical than to be correct.

— Benjamin Disraeli

Page 10: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

executive editoRCarol Long

pRoject editoRVictoria Swider

technical editoRRichard Rowe

pRoduction editoRKathleen Wisor

copy editoRKim Cofer

editoRial manageRMary Beth Wakefield

FReelanceR editoRial manageRRosemarie Graham

aSSociate diRectoR oF maRketingDavid Mayhew

maRketing manageRAshley Zurcher

BuSineSS manageRAmy Knies

pRoduction manageRTim Tate

vice pReSident and executive gRoup puBliSheRRichard Swadley

vice pReSident and executive puBliSheRNeil Edde

aSSociate puBliSheRJim Minatel

pRoject cooRdinatoR, coveRKatie Crocker

compoSitoRCraig Woods, Happenstance Type-O-Rama

pRooFReadeRJames Saturnio, Word One

indexeRJohn Sleeva

coveR deSigneRLeAndra Young

coveR image© iStock / Mark Wragg

cReditS

Page 11: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

aBout the authoR

maRk gaRdeneR (http://www.gardenersown.co.uk) is an ecologist, lecturer, and writer working in the UK. He has a passion for the natu-ral world and for learning new things. Originally he worked in optics, but returned to education in 1996 and eventually gained his doctorate in ecology and evolutionary biology. This work involved a lot of data analysis, and he became interested in R as a tool to help in research. He is currently self-employed and runs courses in ecology, data analy-

sis, and R for a variety of organizations. Mark lives in rural Devon with his wife Christine (a biochemist), and still enjoys the natural world and learning new things.

Page 12: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

acknowledgmentS

FiRSt oF all my thankS go out to the R project team and the many authors and programmers who work tirelessly to make this a peerless program. I would also like to thank my wife, Christine, who has had to put up with me during this entire process, and in many senses became an R-widow! Thanks to Wiley, for asking me to do this book, including Paul Reese, Carol Long, and Victoria Swider. I couldn’t have done it without you. Thanks also to Richard Rowe, the technical reviewer, who first brought my attention to R and its compelling (and rather addictive) power.

Last but not least, thanks to the R community in general. I learned to use R largely by trial and error, and using the vast wealth of knowledge that is in this community. I hope that I have managed to distill this knowledge into a worthy package for future devotees of R.

— Mark Gardener

Page 13: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

contentS

IntroductIon xxi

chapteR 1: intRoducing R: what it iS and how to get it 1

getting the Hang of R 2The R Website 3Downloading and Installing R from CRAN 3

Installing R on Your Windows Computer 4Installing R on Your Macintosh Computer 7Installing R on Your Linux Computer 7

Running the R Program 8Finding Your Way with R 10

Getting Help via the CRAN Website and the Internet 10The Help Command in R 10

Help for Windows Users 11Help for Macintosh Users 11Help for Linux Users 13Help For All Users 13

Anatomy of a Help Item in R 14Command Packages 16

Standard Command Packages 16What Extra Packages Can Do for You 16How to Get Extra Packages of R Commands 18

How to Install Extra Packages for Windows Users 18How to Install Extra Packages for Macintosh Users 18How to Install Extra Packages for Linux Users 19

Running and Manipulating Packages 20Loading Packages 21Windows-Specific Package Commands 21Macintosh-Specific Package Commands 21Removing or Unloading Packages 22

Summary 22

chapteR 2: StaRting out: Becoming FamiliaR with R 25

Some Simple Math 26Use R Like a Calculator 26Storing the Results of Calculations 29

contentS

chapteR 1: cRedit

chapteR 2: aBout the authoR

chapteR 3: acknowledgment

Ion

Who This Book is For

is Structured need to Use This Book

errata

chapteR 4: intRoducing R: what it iS and how to get it

getting the Hang of R

chapteR 5: StaRting out: Becoming FamiliaR with R

getting Data into R

named Objects items

items examining Data Structure

Page 14: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xii

CONTENTS

Reading and getting Data into R 30Using the combine Command for Making Data 30

Entering Numerical Items as Data 30Entering Text Items as Data 31

Using the scan Command for Making Data 32Entering Text as Data 33Using the Clipboard to Make Data 33Reading a File of Data from a Disk 35

Reading Bigger Data Files 37The read .csv() Command 37Alternative Commands for Reading Data in R 39Missing Values in Data Files 40

Viewing named Objects 41Viewing Previously Loaded Named-Objects 42

Viewing All Objects 42Viewing Only Matching Names 42

Removing Objects from R 44Types of Data items 45

Number Data 45Text Items 45Converting Between Number and Text Data 46

The Structure of Data items 47Vector Items 48Data Frames 48Matrix Objects 49List Objects 49

examining Data Structure 49Working with History Commands 51

Using History Files 52Viewing the Previous Command History 52Saving and Recalling Lists of Commands 52Alternative History Commands in Macintosh OS 52

Editing History Files 53Saving Your Work in R 54

Saving the Workspace on Exit 54Saving Data Files to Disk 54

Save Named Objects 54Save Everything 55

Reading Data Files from Disk 56Saving Data to Disk as Text Files 57

Writing Vector Objects to Disk 58Writing Matrix and Data Frame Objects to Disk 58

Page 15: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xiii

CONTENTS

Writing List Objects to Disk 59Converting List Objects to Data Frames 60

Summary 61

chapteR 3: StaRting out: woRking with oBjectS 65

Manipulating Objects 65Manipulating Vectors 66

Selecting and Displaying Parts of a Vector 66Sorting and Rearranging a Vector 68Returning Logical Values from a Vector 70

Manipulating Matrix and Data Frames 70Selecting and Displaying Parts of a Matrix or Data Frame 71Sorting and Rearranging a Matrix or Data Frame 74

Manipulating Lists 76Viewing Objects within Objects 77

Looking Inside Complicated Data Objects 77Opening Complicated Data Objects 78Quick Looks at Complicated Data Objects 80Viewing and Setting Names 82Rotating Data Tables 86

Constructing Data Objects 86Making Lists 87Making Data Frames 88Making Matrix Objects 89Re-ordering Data Frames and Matrix Objects 92

Forms of Data Objects: Testing and Converting 96Testing to See What Type of Object You Have 96Converting from One Object Form to Another 97

Convert a Matrix to a Data Frame 97Convert a Data Frame into a Matrix 98Convert a Data Frame into a List 99Convert a Matrix into a List 100Convert a List to Something Else 100

Summary 104

chapteR 4: data: deScRiptive StatiSticS and taBulation 107

Summary Commands 108Summarizing Samples 110

Summary Statistics for Vectors 110Summary Commands With Single Value Results 110Summary Commands With Multiple Results 113

Page 16: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xiv

CONTENTS

Cumulative Statistics 115Simple Cumulative Commands 115Complex Cumulative Commands 117

Summary Statistics for Data Frames 118Generic Summary Commands for Data Frames 119Special Row and Column Summary Commands 119The apply() Command for Summaries on Rows or Columns 120

Summary Statistics for Matrix Objects 120Summary Statistics for Lists 121

Summary Tables 122Making Contingency Tables 123

Creating Contingency Tables from Vectors 123Creating Contingency Tables from Complicated Data 123Creating Custom Contingency Tables 126Creating Contingency Tables from Matrix Objects 128

Selecting Parts of a Table Object 130Converting an Object into a Table 132Testing for Table Objects 133Complex (Flat) Tables 134

Making “Flat” Contingency Tables 134Making Selective “Flat” Contingency Tables 138

Testing “Flat” Table Objects 139Summary Commands for Tables 139Cross Tabulation 142

Testing Cross-Table (xtabs) Objects 144A Better Class Test 144Recreating Original Data from a Contingency Table 145Switching Class 146

Summary 147

chapteR 5: data: diStRiBution 151

Looking at the Distribution of Data 151Stem and Leaf Plot 152Histograms 154Density Function 158

Using the Density Function to Draw a Graph 159Adding Density Lines to Existing Graphs 160

Types of Data Distribution 161The Normal Distribution 161Other Distributions 164Random Number Generation and Control 166Random Numbers and Sampling 168

Page 17: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xv

CONTENTS

The Shapiro-Wilk Test for Normality 171The Kolmogorov-Smirnov Test 172Quantile-Quantile Plots 174

A Basic Normal Quantile-Quantile Plot 174Adding a Straight Line to a QQ Plot 174Plotting the Distribution of One Sample Against Another 175

Summary 177

chapteR 6: Simple hypotheSiS teSting 181

Using the Student’s t-test 181Two-Sample t-Test with Unequal Variance 182Two-Sample t-Test with Equal Variance 183One-Sample t-Testing 183Using Directional Hypotheses 183Formula Syntax and Subsetting Samples in the t-Test 184

The Wilcoxon U-Test (Mann-Whitney) 188Two-Sample U-Test 189One-Sample U-Test 189Using Directional Hypotheses 189Formula Syntax and Subsetting Samples in the U-test 190

Paired t- and U-Tests 193Correlation and Covariance 196

Simple Correlation 197Covariance 199Significance Testing in Correlation Tests 199Formula Syntax 200

Tests for Association 203Multiple Categories: Chi-Squared Tests 204

Monte Carlo Simulation 205Yates’ Correction for 2 n 2 Tables 206

Single Category: Goodness of Fit Tests 206Summary 210

chapteR 7: intRoduction to gRaphical analySiS 215

Box-whisker Plots 215Basic Boxplots 216Customizing Boxplots 217Horizontal Boxplots 218

Scatter Plots 222Basic Scatter Plots 222Adding Axis Labels 223

Page 18: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xvi

CONTENTS

Plotting Symbols 223Setting Axis Limits 224Using Formula Syntax 225Adding Lines of Best-Fit to Scatter Plots 225

Pairs Plots (Multiple Correlation Plots) 229Line Charts 232

Line Charts Using Numeric Data 232Line Charts Using Categorical Data 233

Pie Charts 236Cleveland Dot Charts 239Bar Charts 245

Single-Category Bar Charts 245Multiple Category Bar Charts 250

Stacked Bar Charts 250Grouped Bar Charts 250

Horizontal Bars 253Bar Charts from Summary Data 253

Copy graphics to Other Applications 256Use Copy/Paste to Copy Graphs 257Save a Graphic to Disk 257

Windows 257Macintosh 258Linux 258

Summary 259

chapteR 8: FoRmula notation and complex StatiSticS 263

examples of Using Formula Syntax for Basic Tests 264Formula notation in graphics 266Analysis of Variance (AnOVA) 268

One-Way ANOVA 268Stacking the Data before Running Analysis of Variance 269Running aov() Commands 270

Simple Post-hoc Testing 271Extracting Means from aov() Models 271Two-Way ANOVA 273

More about Post-hoc Testing 275Graphical Summary of ANOVA 277Graphical Summary of Post-hoc Testing 278

Extracting Means and Summary Statistics 281Model Tables 281Table Commands 283

Page 19: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xvii

CONTENTS

Interaction Plots 283More Complex ANOVA Models 289Other Options for aov() 290

Replications and Balance 290Summary 292

chapteR 9: manipulating data and extRacting componentS 295

Creating Data for Complex Analysis 295Data Frames 296Matrix Objects 299Creating and Setting Factor Data 300Making Replicate Treatment Factors 304Adding Rows or Columns 306

Summarizing Data 312Simple Column and Row Summaries 312Complex Summary Functions 313

The rowsum() Command 314The apply() Command 315Using tapply() to Summarize Using a Grouping Variable 316The aggregate() Command 319

Summary 323

chapteR 10: RegReSSion (lineaR modeling) 327

Simple Linear Regression 328Linear Model Results Objects 329

Coefficients 330Fitted Values 330Residuals 330Formula 331Best-Fit Line 331

Similarity between lm() and aov() 334Multiple Regression 335

Formulae and Linear Models 335Model Building 337

Adding Terms with Forward Stepwise Regression 337Removing Terms with Backwards Deletion 339Comparing Models 341

Curvilinear Regression 343Logarithmic Regression 344Polynomial Regression 345

Page 20: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xviii

CONTENTS

Plotting Linear Models and Curve Fitting 347Best-Fit Lines 348

Adding Line of Best-Fit with abline() 348Calculating Lines with fitted() 348Producing Smooth Curves using spline() 350

Confidence Intervals on Fitted Lines 351Summarizing Regression Models 356

Diagnostic Plots 356Summary of Fit 357

Summary 359

chapteR 11: moRe aBout gRaphS 363

Adding elements to existing Plots 364Error Bars 364

Using the segments() Command for Error Bars 364Using the arrows() Command to Add Error Bars 368

Adding Legends to Graphs 368Color Palettes 370Placing a Legend on an Existing Plot 371

Adding Text to Graphs 372Making Superscript and Subscript Axis Titles 373Orienting the Axis Labels 375Making Extra Space in the Margin for Labels 375Setting Text and Label Sizes 375Adding Text to the Plot Area 376Adding Text in the Plot Margins 378Creating Mathematical Expressions 379

Adding Points to an Existing Graph 382Adding Various Sorts of Lines to Graphs 386

Adding Straight Lines as Gridlines or Best-Fit Lines 386Making Curved Lines to Add to Graphs 388Plotting Mathematical Expressions 390Adding Short Segments of Lines to an Existing Plot 393Adding Arrows to an Existing Graph 394

Matrix Plots (Multiple Series on One graph) 396Multiple Plots in One Window 399

Splitting the Plot Window into Equal Sections 399Splitting the Plot Window into Unequal Sections 402

exporting graphs 405Using Copy and Paste to Move a Graph 406Saving a Graph to a File 406

Page 21: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xix

CONTENTS

Windows 406Macintosh 406Linux 406

Using the Device Driver to Save a Graph to Disk 407PNG Device Driver 407PDF Device Driver 407Copying a Graph from Screen to Disk File 408Making a New Graph Directly to a Disk File 408

Summary 410

chapteR 12: wRiting youR own ScRiptS: Beginning to pRogRam 415

Copy and Paste Scripts 416Make Your Own Help File as Plaintext 416Using Annotations with the # Character 417

Creating Simple Functions 417One-Line Functions 417Using Default Values in Functions 418Simple Customized Functions with Multiple Lines 419Storing Customized Functions 420

Making Source Code 421Displaying the Results of Customized Functions and Scripts 421Displaying Messages as Part of Script Output 422

Simple Screen Text 422Display a Message and Wait for User Intervention 424

Summary 428

appendix: anSweRS to exeRciSeS 433

Index 461

Page 22: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how
Page 23: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

intRoduction

thiS Book iS aBout data analySiS and the programming language called R. This is rapidly becoming the de facto standard among professionals, and is used in every conceivable discipline from science and medicine to business and engineering.

R is more than just a computer program; it is a statistical programming environment and language. R is free and open source and is therefore available to everyone with a computer. It is very powerful and flexible, but it is also unlike most of the computer programs you are likely used to. You have to type commands directly into the program to make it work for you. Because of this, and its complexity, R can be hard to get a grip on.

This book delves into the language of R and makes it accessible using simple data examples to explore its power and versatility. In learning how to “speak R,” you will unlock its potential and gain better insights into tackling even the most complex of data analysis tasks.

who thiS Book iS FoR

This book is for anyone who needs to analyze any data, whatever their discipline or line of work. Whether you are in science, business, medicine, or engineering, you will have data to analyze and results to present. R is powerful and flexible and completely cross-platform. This means you can share data and results with anyone. R is backed by a huge project team, so being free does not mean being inferior!

If you are completely new to R, this book will enable you to get it and start to become familiar with it. There is no assumption that you know anything about the program to begin with. If you are already familiar with R, you will find this book a useful reference that you can call upon time and time again; the first chapter is largely concerned with installing R, so you may want to skip to Chapter 2.

This book is not about statistical analyses, so some familiarity with basic analytical methods is helpful (but not obligatory). The book deals with the means to make R work for you; this means learning the language of R rather than learning statistics. Once you are familiar with R you will be empowered to use it to undertake a huge variety of analytical tasks, more than can be conveniently packaged into a single book. R also produces presentation-quality graphics and this book leads you through the complexities of that.

what thiS Book coveRS

R is a computer program and statistical programming language/environment. It allows a wide range of analytical methods to be used and produces presentation-quality graphics. This book covers the language of R, and leads you toward a better understanding of how to get R to do the things you need. There is less emphasis on the actual statistical tests; indeed, R is so flexible that the list of tests

Page 24: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xxii

introduction

it can perform is far too large to be covered in an introductory book such as this. Rather, the aim is to become familiar with the language of R and to carry out some of the more commonly used statistical methods. In this way, you can strike out on your own and explore the full potential of R for yourself.

So, the focus is on the operation of R itself. Along the way you learn how to carry out a range of commonly used statistical methods, including analysis of variance (ANOVA) and linear regression, which are widely used in many fields and, therefore, important to know. You also learn a range of ways to produce a wide variety of graphics that should suit your needs.

This book covers most recent versions of R. The R program does change from time to time as new versions are released. However, most of the commands you will need to know have not changed, and even older (in computer terms) versions will work quite happily.

how thiS Book iS StRuctuRed

The book has a general progressive character, and later chapters tend to build on skills you learned earlier. Therefore if you are a beginner, you will probably find it most useful to start at the beginning and work your way through in a progressive manner. If you are a more seasoned user, you may want to use selected chapters as reference material, to refresh your skills.

No approach to learning R is universally adequate, but I have tried to provide the most logical path possible. For example, learning to produce graphics is very important, but unless you know what kinds of analyses you are likely to need to represent, making these graphs might seem a bit prosaic. Therefore, the main graphics chapter appears after some of the chapters on analysis.

In general terms, the book begins with notes on how to get and install R, and how to access the help system. Next you are introduced to the basics of data—how to get data into R, for example. After this you find out how to manipulate data, carry out some basic statistical analyses, and begin to tackle graphics. Later you learn some more advanced analytical methods and return to graphics. Finally, you look at ways to use R to create your own programs.

Each chapter begins with an overview of the topics you will learn. The text contains many examples and is written in a “copy me” style. Throughout the text, all the concepts are illustrated with simple examples. You can download the data from the companion website and follow along as you read (details on this are discussed shortly). The book contains a variety of activities that you are urged to follow; each is designed to help you with an important topic. The chapters all end with a series of exercises that help you to consolidate your learning (the solutions are in the appendix). Finally, the chapters end with a brief summary of what you learned and a table illustrating the topics and some key points, which are useful as reference material. Following is a brief description of each chapter.

Chapter 1: Introducing R: What It Is and How to Get It—In this chapter you see how to get R and install it on your computer. You also learn how to access the built-in help system and find out about additional packages of useful analytical routines that you can add to R.

Page 25: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xxiii

introduction

Chapter 2: Starting Out: Becoming Familiar with R—This chapter builds some familiarity with working with R, beginning with some simple math and culminating in importing and making data objects that you can work with (and saving data to disk for later use).

Chapter 3: Starting Out: Working With Objects—This chapter deals with manipulating the data that you have created or imported. These are important tasks that underpin many of the later exercises. The skills you learn here will be put to use over and over again.

Chapter 4: Data: Descriptive Statistics and Tabulation—This chapter is all about summariz-ing data. Here you learn about basic summary methods, including cumulative statistics. You also learn how about cross-tabulation and how to create summary tables.

Chapter 5: Data: Distribution—In this chapter you look at visualizing data using graphi-cal methods—for example, histograms—as well as mathematical ones. This chapter also includes some notes about random numbers and different types of distribution (for example, normal and Poisson).

Chapter 6: Simple Hypothesis Testing—In this chapter you learn how to carry out some basic statistical methods such as the t-test, correlation, and tests of association. Learning how to do these is helpful for when you have to carry out more complex analyses and also illustrates a range of techniques for using R.

Chapter 7: Introduction to Graphical Analysis—In this chapter you learn how to produce a range of graphs including bar charts, scatter plots, and pie charts. This is a “first look” at making graphs, but you return to this subject in Chapter 11, where you learn how to turn your graphs from merely adequate to stunning.

Chapter 8: Formula Notation and Complex Statistics—As your analyses become more complex, you need a more complex way to tell R what you want to do. This chapter is con-cerned with an important element of R: how to define complex situations. The chapter has two main parts. The first part shows how the formula notation can be used with simple situations. The second part uses an important analytical method, analysis of variance, as an illustration. The rest of the chapter is devoted to ANOVA. This is an important chapter because the ability to define complex analytical situations is something you will inevitably require at some point.

Chapter 9: Manipulating Data and Extracting Components—This chapter builds on the previous one. Now that you have seen how to define more complex analytical situations, you learn how to make and rearrange your data so that it can be analyzed more easily. This also builds on knowledge gained in Chapter 3. In many cases, when you have carried out an analysis you will need to extract data for certain groups; this chapter also deals with that, giving you more tools that you will need to carry out complex analyses easily.

Chapter 10: Regression (Linear Modeling)—This chapter is all about regression. It builds on earlier chapters and covers various aspects of this important analytical method. You learn how to carry out basic regression, as well as complex model building and curvilinear regres-sion. It is also important because it illustrates some useful aspects of R (for example, how to dissect results). The later parts of the chapter deal with graphical aspects of regression, such as how to add lines of best fit and confidence intervals.

Page 26: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xxiv

introduction

Chapter 11: More About Graphs—This chapter builds on the earlier chapter on graphics (Chapter 7) and also from the previous chapter on regression. It shows you how to produce more customized graphs from your data. For example, you learn how to add text to plots and axes, and how to make superscript and subscript text and mathematical symbols. You learn how to add legends to plots, and how to add error bars to bar charts or scatter plots. Finally, you learn how to export graphs to disk as high-quality graphics files, suitable for publication.

Chapter 12: Writing Your Own Scripts: Beginning to Program—In this chapter you learn how to start producing customized functions and simple scripts that can automate your workflow, and make complex and repetitive tasks a lot easier.

what you need to uSe thiS Book

The only things you need to use this book are a computer and enthusiasm! The R program works on any operating system, so you can use Windows, Macintosh, or Linux (any version). R even works quite adequately on ancient (in computer terms) computers, so you do not need anything particularly hi-spec. An Internet connection is required at some point because you need to get R from the R-project website. However, it is perfectly possible to download the installation files onto a separate computer and transfer them to your working machine.

If you already have a version of R, it is not necessary to get the latest version. R is continually changing and improving, but the older versions of R will most likely work with this book because the basic command set has changed relatively little. Having said that, I suggest you update your version of R if it is older than 2009.

conventionS

To help you get the most from the text and keep track of what’s happening, we’ve used a number of conventions throughout the book.

The commands you need to type into R and the output you get from R are shown in a monospace font. Each example that shows lines that are typed by the user begins with the > symbol, which mimics the R cursor like so:

> help()

Lines that begin with something other than the > symbol represent the output from R (but look out for typed lines that are long and spread over more than one line), so in the following example the first line was typed by the user and the second line is the result:

> data1 [1] 3 5 7 5 3 2 6 8 5 6 9

Page 27: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xxv

introduction

tRy it out

The Try It Out is an exercise you should work through, following the text in the book.

1 . They usually consist of a set of steps.

2 . Each step has a number.

3 . Follow the steps through with your copy of the database.

How It Works

After each Try It Out, the code you’ve typed is explained in detail.

WARNING Boxes with a warning icon like this one hold important, not-to-be forgotten information that is directly relevant to the surrounding text.

NOTE The pencil icon indicates notes, tips, hints, tricks, and asides to the current discussion.

As for styles in the text:

➤➤ We highlight new terms and important words when we introduce them.

➤➤ We show keyboard strokes like this: Ctrl+A.

➤➤ We show filenames, URLs, and code within the text like so: persistence.properties.

➤➤ We present code in two different ways:

We use a monofont type with no highlighting for most code examples.We use bold to emphasize code that’s particularly important in the present context.

SouRce code

As you work through the examples in this book, you may choose either to type in all the data and code manually or to use the source code and data object files that accompany the book. All of the data and source code used in this book is available for download at http://www.wrox.com. You will find the data sets that you need for each example activity are accompanied by a download icon and note indicating the name of the data file so you know it’s available for download and can eas-ily locate it in the download file. Once at the site, simply locate the book’s title (either by using the Search box or by using one of the title lists) and click the Download Code link on the book’s detail page to obtain all the source code for the book.

Page 28: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xxvi

introduction

There will only be one file to download and it is called Beginning.RData. This one file contains all the example datasets and scripts you need for the whole book. Once you have the file on your com-puter you can load it into R by one of several methods:

➤➤ For Windows or Mac you can drag the Beginning.RData file icon onto the R program icon; this will open R if it is not already running and load the data. If R is already open, the data will be appended to anything you already have in R; otherwise only the data in the file will be loaded.

➤➤ If you have Windows or Macintosh you can also load the file using menu commands or use a command typed into R:

➤➤ For Windows use File d Load Workspace, or type the following command in R:

> load(file.choose())

➤➤ For Mac use Workspace d Load Workspace File, or type the following command in R (same as in Windows):

> load(file.choose())

➤➤ If you have Linux then you can use the load() command but must specify the filename (in quotes) exactly, for example:

> load(“Beginning.RData”)

The Beginning.RData file must be in your default working directory and if it is not you must specify the location as part of the filename.

NOTE Because many books have similar titles, you may find it easiest to search by ISBN; this book’s ISBN is 978-1-118-164303.

Alternatively, you can go to the main Wrox code download page at http://www.wrox.com/dynamic/books/download.aspx to see the code available for this book and all other Wrox books.

eRRata

We make every effort to ensure that there are no errors in the text or in the code. However, no one is perfect, and mistakes do occur. If you find an error in one of our books, like a spelling mistake or faulty piece of code, we would be very grateful for your feedback. By sending in errata you may save another reader hours of frustration and at the same time you will be helping us provide even higher quality information.

To find the errata page for this book, go to http://www.wrox.com and locate the title using the Search box or one of the title lists. Then, on the book details page, click the Book Errata link. On this page you can view all errata that has been submitted for this book and posted by Wrox editors.

Page 29: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how

xxvii

introduction

NOTE A complete book list including links to each book’s errata is also available at www.wrox.com/misc-pages/booklist.shtml.

If you don’t spot “your” error on the Book Errata page, go to www.wrox.com/contact/techsupport .shtml and complete the form there to send us the error you have found. We’ll check the information and, if appropriate, post a message to the book’s errata page and fix the problem in subsequent editions of the book.

p2p .wRox .com

For author and peer discussion, join the P2P forums at p2p.wrox.com. The forums are a web-based system for you to post messages relating to Wrox books and related technologies and interact with other readers and technology users. The forums offer a subscription feature to e-mail you topics of interest of your choosing when new posts are made to the forums. Wrox authors, editors, other industry experts, and your fellow readers are present on these forums.

At http://p2p.wrox.com you will find a number of different forums that will help you not only as you read this book, but also as you develop your own applications. To join the forums, just follow these steps:

1 . Go to p2p.wrox.com and click the Register link.

2 . Read the terms of use and click Agree.

3 . Complete the required information to join as well as any optional information you wish to provide and click Submit.

4 . You will receive an e-mail with information describing how to verify your account and complete the joining process.

NOTE You can read messages in the forums without joining P2P but in order to post your own messages, you must join.

Once you join, you can post new messages and respond to messages other users post. You can read messages at any time on the web. If you would like to have new messages from a particular forum e-mailed to you, click the Subscribe to this Forum icon by the forum name in the forum listing.

For more information about how to use the Wrox P2P, be sure to read the P2P FAQs for answers to questions about how the forum software works as well as many common questions specific to P2P and Wrox books. To read the FAQs, click the FAQ link on any P2P page.

Page 30: Beginning R - download.e-bookshelf.dedownload.e-bookshelf.de/.../0000/6294/42/L-G-0000629442-00023650… · content S IntroductIon xxi chapteR 1: intRoducing R: what it iS and how