7 Samples and Populations (Ch. 7)
Drennan dwells at some length on sampling because it is so fundamental. It’s important to remember that often our most basic concern with sampling is understanding the nature of the sample we have, and how it relates to the population in which we are interested - but a much more immediate and practical concern is how to efficiently draw random samples. A simple but very powerful tool for doing this in R is sample(), which can replace the table of random numbers that Drennan discusses.
Suppose that you had gridded an area in 1 ha blocks and wanted to select a random sample of grid squares in which to collect cultural material, or crop varieties, or butterflies. If the area you were studying was 10 km x 7.5 km, you would have 7500 ha to choose from [(10000 * 7500)/10000], from which you might want to sample only 100. Perhaps the simplest way to sample these would be to assign a number to each hectare block within the grid, and then randomly select 100.
grids <- sample(1:7500, 100) #randomly select 100 numbers from 1-7500, w/o replacement
#if you have reason to sample with replacement instead, use replace = TRUE in sample()
grids #by default these are ordered as selected## [1] 32 4352 464 5601 6022 5396 1115 7036 5793 1199 4532 3191 861 7151 471
## [16] 4753 937 419 6818 4507 1492 1279 334 2177 7305 6482 7495 5790 4947 6516
## [31] 3670 1415 3659 1007 5199 1812 2884 4373 6872 3725 6345 6254 217 4056 2186
## [46] 3633 2794 7296 1266 2669 3600 2219 4080 2282 7149 1630 7259 557 4657 3097
## [61] 3953 1609 1060 449 6791 6448 6195 5425 1211 1750 2479 4436 2052 4884 1338
## [76] 2211 2119 3040 5991 6649 5669 3368 5129 5016 5076 6422 2836 2713 4375 3929
## [91] 3625 2175 4139 6548 4157 1924 4445 7181 4236 7123
Probably - think about the practical realities of working your way around the landscape collecting samples - you’d find it more practical to order these sequentially.
## [1] 32 217 334 419 449 464 471 557 861 937 1007 1060 1115 1199 1211
## [16] 1266 1279 1338 1415 1492 1609 1630 1750 1812 1924 2052 2119 2175 2177 2186
## [31] 2211 2219 2282 2479 2669 2713 2794 2836 2884 3040 3097 3191 3368 3600 3625
## [46] 3633 3659 3670 3725 3929 3953 4056 4080 4139 4157 4236 4352 4373 4375 4436
## [61] 4445 4507 4532 4657 4753 4884 4947 5016 5076 5129 5199 5396 5425 5601 5669
## [76] 5790 5793 5991 6022 6195 6254 6345 6422 6448 6482 6516 6548 6649 6791 6818
## [91] 6872 7036 7123 7149 7151 7181 7259 7296 7305 7495
As Drennan (p85) is at pains to point out - and as you would be very wise to note! - random is not the same thing as representative. The concepts underlying Drennan’s discussion of random and nonrandom samples, sampling bias, and representativeness are at least as important as any statistical technique. As Drennan notes (p88), these are issues fundamental to inference generally, not just to statistical inference (i.e., you cannot avoid the need to think about these issues by simply eschewing statistics!).
Note that R code that generates random samples is very unlikely to draw the same number twice! So…if you run the code a second time, you’re likely to get a different answer. If you want to, for illustrative purposes, makes sure that you always draw the same random number, use the set.seed() function.
Compare the result of running both lines of code below two or three times to the result of running only the second line of code two or three times.