The Discipline II: Fieldwork and the Fieldwork Tradition, the Methods from Ethnography and Observation to Grounded Theory, Excavation and GIS, Statistics from Variables and Sampling to Regression, and Content, Discourse and Narrative Analysis

The second half of Unit I is the toolkit. Fieldwork is what made anthropology a discipline rather than a library subject, and the syllabus lists its methods by name: ethnography, observation, interview, case study, life history, focus group, PRA and RRA, the genealogical method, schedules and questionnaires, grounded theory, exploration and excavation, and GIS. It then asks for the statistics an anthropologist actually uses — variables, sampling, central tendency and dispersion, parametric and non-parametric tests, bivariate and multivariate analysis including linear and logistic regression — and for three techniques of analysing text and talk: content, discourse and narrative analysis. This chapter takes each in turn, names the person or the study each method is attached to, and works a mean, a standard deviation and a chi-square, because they are asked as numbers to be computed as well as definitions to be recalled.

1. Fieldwork and the fieldwork tradition

Fieldwork is the prolonged, first-hand study of a people in their own setting, in their own language, through daily participation in their life. The armchair evolutionists of the 1870s (Tylor, Frazer, Morgan, who did visit the Iroquois) worked from correspondents' reports; the shift to the field came in stages. Boas wintered on Baffin Island among the Inuit in 1883–84. The Cambridge expedition to the Torres Straits in 1898 (Haddon, Rivers, Seligman, Myers, McDougall) took a team of scientists to collect systematically, and Rivers developed the genealogical method there. Radcliffe-Brown worked in the Andamans in 1906–08 and Seligman among the Vedda. Then Malinowski, marooned by the First World War, spent about two years in the Trobriand Islands (1915–16, 1917–18), living in a tent in the village, learning Kiriwina and writing Argonauts of the Western Pacific (1922), whose introduction laid down the method: pitch the tent among the natives, record the "imponderabilia of actual life", and grasp "the native's point of view". Participant observation, a year or more in one place, a field diary and a monograph became the standard. In India the tradition began with S.C. Roy's Chotanagpur monographs, continued with Majumdar among the Ho and Elwin among the Baiga and Muria, and matured in the village studies of the 1950s (Srinivas at Rampura, Dube at Shamirpet). Its ethical side — informed consent, confidentiality, no harm, reciprocity to the community — is now codified by the professional associations, and reflexivity (attention to the fieldworker's own position) is expected of every ethnography.

📖 Malinowski's diary
A Diary in the Strict Sense of the Term, published in 1967 after his death, showed the founder of participant observation bored, ill and contemptuous of his hosts in private. The shock it caused made reflexivity — writing the fieldworker's own feelings and position into the account — a requirement rather than a confession, and it is the usual starting point of questions on the "crisis of representation" (Unit VIII).

2. The methods, one by one

MethodWhat it isAttached name or study
EthnographyHolistic description of a people from long residence; also the written monographMalinowski (Argonauts, 1922); Geertz's thick description; multi-sited ethnography (Marcus, 1995)
ObservationParticipant (the researcher joins the activity) or non-participant; controlled or uncontrolled; overt or covertMalinowski; the Chicago school; Gold's four roles (complete participant to complete observer, 1958)
InterviewStructured (fixed questions), semi-structured (guide), unstructured (conversation); key-informant interviews with knowledgeable insidersKey informants: Boas's George Hunt among the Kwakiutl; Marriott's and Srinivas's village informants
Case studyIntensive study of one unit (a person, family, dispute, village) in its context; the extended-case method follows one dispute through timeGluckman's Manchester school; van Velsen's situational analysis (1967); Turner's social drama among the Ndembu
Life historyA person's life narrated at length, as a window on culture and changeRadin's Crashing Thunder (1926); Oscar Lewis's The Children of Sánchez (1961); Mandelbaum's "The Study of Life History: Gandhi" (1973)
Focus groupA guided discussion among six to twelve people on a set topic, recorded and analysed for the interaction as much as the contentMerton's focused interview (1940s); Krueger; widely used in health and development research
RRA and PRARapid Rural Appraisal (1980s): quick multidisciplinary field appraisal by outsiders; Participatory Rural Appraisal (1990s): villagers themselves map, rank and diagram — social and resource maps, transect walks, seasonal calendars, wealth ranking, Venn diagramsRobert Chambers (Rural Development: Putting the Last First, 1983; PRA papers 1994); the Khon Kaen RRA conference, 1985
Genealogical methodRecording kin of each informant by a fixed set of terms (father, mother, child, husband, wife) to reconstruct descent, marriage, residence and the terminologyW.H.R. Rivers, Torres Straits 1898; "The Genealogical Method of Anthropological Inquiry" (1910); Todas (1906)
Schedule and questionnaireBoth are lists of questions; a schedule is filled by the investigator face to face (usable with non-literate respondents, higher response, costlier), a questionnaire by the respondent (mailed or online, cheaper, lower response)The Census house-list schedule; Majumdar's and Dube's village surveys
Grounded theoryTheory built inductively from data by constant comparison: open, axial and selective coding, memos, theoretical sampling until saturationGlaser and Strauss, The Discovery of Grounded Theory (1967); Strauss and Corbin (1990); Charmaz's constructivist version
Exploration and excavationExploration: surface survey, village-to-village and transect survey, aerial and satellite imagery, trial trenches; excavation: vertical (deep, for sequence) or horizontal (wide, for layout); the grid method; stratigraphy recorded layer by layer; the Harris matrixMortimer Wheeler's grid (Archaeology from the Earth, 1954); Sankalia's vertical digs; Dhavalikar's horizontal exposure of Inamgaon
GISGeographic Information Systems: layered spatial data (sites, soils, rivers, settlements) stored, queried and mapped; with remote sensing it locates sites, models catchments and maps caste or disease distributionsTomlinson (Canada GIS, 1960s); used in Indian archaeology for the Ghaggar–Hakra palaeochannels and in demography for health mapping
⚠️ Schedule or questionnaire; RRA or PRA
Two pairs the paper likes to swap. The schedule is filled by the investigator, the questionnaire by the respondent — so the schedule suits the non-literate and the questionnaire is cheaper. RRA is the outsider's rapid appraisal; PRA hands the pen to the villagers ("handing over the stick", in Chambers's phrase). An option calling PRA "rapid" or the questionnaire "interviewer-administered" is the distractor.

3. Statistics I: variables, sampling, central tendency and dispersion, with a worked mean and standard deviation

A variable is any characteristic that takes different values across cases: the independent variable is the presumed cause, the dependent the effect, and an intervening or confounding variable can sit between or behind them. Stevens's four levels of measurement (1946) decide which statistic is allowed: nominal (categories — caste, blood group; mode, chi-square), ordinal (ranked — wealth rank, Likert scale; median, Spearman's rho), interval (equal units, no true zero — temperature in °C; mean, SD, t-test) and ratio (true zero — height, income; all of these plus ratios). Sampling draws a manageable part of the universe. Probability samples give every unit a known chance of selection: simple random (lottery or random numbers), systematic (every kth house from a random start), stratified (divide by caste or age, then sample each stratum; proportionate or not), cluster and multi-stage (sample villages, then households). Non-probability samples do not: purposive or judgement (the elders who know the ritual), quota, convenience, and snowball (each informant names the next — the method for hidden populations). Sampling error shrinks with sample size; non-sampling error (bad questions, non-response) does not.

The measures of central tendency are the mean (sum divided by n), the median (the middle value when ranked; unaffected by extreme values, so preferred for income) and the mode (the most frequent value; the only one usable for nominal data). In a symmetrical distribution they coincide; in a positively skewed one (a long right tail, as with income) mean > median > mode. The measures of dispersion are the range, the quartile deviation, the mean deviation, the variance (the mean of squared deviations from the mean) and its square root, the standard deviation; the coefficient of variation, CV = (SD ÷ mean) × 100, compares the spread of series in different units. Worked: the heights of five adult men are 160, 165, 170, 175 and 180 cm. Mean = 850 ÷ 5 = 170. Deviations −10, −5, 0, 5, 10; squares 100, 25, 0, 25, 100; sum 250. Population variance = 250 ÷ 5 = 50 and SD = √50 ≈ 7.07 cm; the sample (n − 1) variance is 250 ÷ 4 = 62.5 and SD ≈ 7.91 cm. CV = 7.07 ÷ 170 × 100 ≈ 4.2%. When a question describes the values as the whole group, use the population form (divide by n); when it calls them a sample, divide by n − 1.

🧠 Which average for which scale
Nominal → mode only; ordinal → median (and mode); interval and ratio → mean (and the rest). "The mean caste of the village" is meaningless; "the median landholding" is the right average for a skewed distribution. The paper's favourite distractor is the mean applied to ranked data.

4. Statistics II: parametric and non-parametric tests, bivariate and multivariate analysis, linear and logistic regression, with a worked chi-square

Parametric tests assume interval or ratio data drawn from a normally distributed population with similar variances: the t-test (Gosset, "Student", 1908; one-sample, independent-samples, paired) compares means, the F-test and analysis of variance (Fisher) compare more than two means, and Pearson's product-moment correlation r measures a linear relation between −1 and +1. Non-parametric ("distribution-free") tests need only nominal or ordinal data: the chi-square test (Pearson, 1900) for association between categorical variables, the Mann–Whitney U and Wilcoxon tests in place of the t-test, the Kruskal–Wallis test in place of ANOVA, the sign and run tests, and Spearman's rank correlation rho. Bivariate analysis relates two variables; multivariate analysis relates three or more at once — multiple regression, factor analysis, cluster analysis, discriminant analysis. Simple linear regression fits y = a + bx, with b the slope (change in y per unit x) and a the intercept, and reads a continuous dependent variable — stature from femur length is the forensic classic. Logistic regression is for a binary dependent variable (stunted or not, migrated or not): it models the log of the odds, ln[p/(1 − p)] = a + bx, and each coefficient is read as an odds ratio, e^b, the factor by which the odds of the outcome change per unit of the predictor.

Worked chi-square. In a village, 120 households are classified by whether the head is literate and whether the household uses the public health centre. Observed: literate and uses, 45; literate and does not, 15; non-literate and uses, 30; non-literate and does not, 30. Row totals 60 and 60; column totals 75 and 45. Expected count = (row total × column total) ÷ N: literate-uses 60 × 75 ÷ 120 = 37.5; literate-not 22.5; non-literate-uses 37.5; non-literate-not 22.5. χ² = Σ(O − E)²/E = (7.5²/37.5) + (7.5²/22.5) + (7.5²/37.5) + (7.5²/22.5) = 1.5 + 2.5 + 1.5 + 2.5 = 8.0. Degrees of freedom = (rows − 1)(columns − 1) = 1; the critical value at 0.05 is 3.84 and at 0.01 is 6.63, so the association between literacy and use of the centre is significant at both levels. A goodness-of-fit chi-square works the same way against theoretical proportions: 120 offspring expected in a 3:1 ratio (90 : 30) but observed 75 : 45 give χ² = (15²/90) + (15²/30) = 2.5 + 7.5 = 10 on 1 df, and the 3:1 hypothesis is rejected.

QuestionParametric testNon-parametric equivalent
Do two independent groups differ?Independent-samples t-testMann–Whitney U
Do paired measurements differ (before and after)?Paired t-testWilcoxon signed-rank; sign test
Do three or more groups differ?One-way ANOVA (F)Kruskal–Wallis H
Are two variables related?Pearson's rSpearman's rho (ranks); chi-square (categories)
Can y be predicted from x?Linear regression (continuous y)Logistic regression (binary y; odds ratios)

5. Techniques of analysis: content, discourse and narrative

Content analysis, defined by Berelson (1952) as the objective, systematic and quantitative description of the manifest content of communication, counts: how often a newspaper names a tribe, which adjectives a textbook attaches to "primitive", how many folk tales feature a trickster. Its steps are to define the universe of texts, draw a sample, fix the unit (word, theme, item), build a coding frame, code with two coders for inter-coder reliability, and tabulate; Krippendorff's alpha measures that reliability, and qualitative content analysis codes latent as well as manifest meaning. Discourse analysis treats language as social action: it asks how a way of talking (about "tribal backwardness", about "development") constructs its object, who may speak, and what is made unsayable. Its lineage runs from Foucault's discourses as regimes of knowledge and power, through Fairclough's critical discourse analysis of text, discursive practice and social practice, to conversation analysis (Sacks, Schegloff) of turn-taking in talk. Narrative analysis takes the story itself as the unit: Labov's structure of oral narrative (abstract, orientation, complicating action, evaluation, result, coda), Riessman's thematic, structural, dialogic and visual approaches, and the anthropological use of life stories, myths and illness narratives (Kleinman's illness narratives) to see how people make experience meaningful and how a community's stories order its past.

⚠️ Manifest, latent, discursive
Content analysis in Berelson's sense counts manifest content; discourse analysis reads how talk produces power and subjects; narrative analysis keeps the story whole. An option that makes content analysis "interpretive study of hidden meaning" or discourse analysis "frequency counting" has swapped them.

Key takeaways

  • Fieldwork became the discipline's method through Boas (Baffin Island 1883–84), the Torres Straits expedition (1898), Radcliffe-Brown in the Andamans (1906–08) and above all Malinowski in the Trobriands (1915–18), whose Argonauts (1922) codified participant observation.
  • Names to attach: Rivers to the genealogical method (1910), Glaser and Strauss to grounded theory (1967), Chambers to PRA (1990s) after RRA (1980s), Merton to the focused interview, Gluckman and van Velsen to the extended case, Wheeler to the excavation grid; the schedule is filled by the investigator, the questionnaire by the respondent.
  • Stevens's four scales (nominal, ordinal, interval, ratio) decide the statistic: mode, median, mean; probability samples (random, systematic, stratified, cluster) against non-probability (purposive, quota, convenience, snowball).
  • Worked figures: heights 160–180 cm in steps of 5 give mean 170, population SD ≈ 7.07 (sample SD ≈ 7.91), CV ≈ 4.2%; a 2 × 2 table with expected 37.5/22.5 and deviations of 7.5 gives χ² = 8.0 on 1 df, significant at 0.05 (3.84) and 0.01 (6.63).
  • Parametric tests (t, F/ANOVA, Pearson's r) assume normal interval data; their non-parametric partners are Mann–Whitney, Wilcoxon, Kruskal–Wallis, chi-square and Spearman; linear regression predicts a continuous y, logistic regression a binary y through odds ratios; content analysis counts manifest content (Berelson), discourse analysis reads talk as power (Foucault, Fairclough), narrative analysis keeps the story whole (Labov, Riessman).

Practice questions (10)

Attempt each one before opening the answer. Every explanation names the tempting wrong option as well as the right one, because that is where marks are lost.

  1. The genealogical method of anthropological inquiry, developed in the Torres Straits in 1898, is associated with

    1. A.C. Haddon
    2. Bronislaw Malinowski
    3. W.H.R. Rivers
    4. Franz Boas
    Show answer

    Answer: C — W.H.R. Rivers

    Rivers devised the method on the Torres Straits expedition, applied it among the Todas (1906) and set it out in a 1910 paper: kin are recorded through a fixed set of terms so that descent, marriage and terminology can be reconstructed. Haddon led the expedition.
  2. The heights of five men are 160, 165, 170, 175 and 180 cm. What is the population variance of these heights, in cm²?

    Numerical answer — type the value.

    Show answer

    Answer: 50

    The mean is 170; the squared deviations are 100, 25, 0, 25 and 100, summing to 250; dividing by n = 5 gives a population variance of 50 (the sample form, dividing by 4, would give 62.5). The standard deviation is √50 ≈ 7.07 cm.
  3. Which pair correctly distinguishes a schedule from a questionnaire?

    1. A schedule is mailed to respondents; a questionnaire is administered face to face
    2. A schedule is used for literate respondents and a questionnaire for non-literate ones
    3. A schedule has only open questions; a questionnaire only closed ones
    4. A schedule is filled in by the investigator in the respondent's presence; a questionnaire is filled in by the respondent
    Show answer

    Answer: D — A schedule is filled in by the investigator in the respondent's presence; a questionnaire is filled in by the respondent

    Both are lists of questions; the difference is who writes the answers. Because the investigator fills the schedule, it works with non-literate respondents and gets a higher response rate, at a higher cost; the self-completed questionnaire is cheaper and needs literacy.
  4. In a 2 × 2 contingency table of 120 households, the observed frequencies exceed or fall short of the expected frequencies (37.5, 22.5, 37.5, 22.5) by 7.5 in every cell. The chi-square value is

    1. 30.0, significant at the 0.01 level
    2. 8.0, significant at the 0.05 level on 1 degree of freedom
    3. 2.0, not significant on 1 degree of freedom
    4. 8.0, tested on 3 degrees of freedom
    Show answer

    Answer: B — 8.0, significant at the 0.05 level on 1 degree of freedom

    χ² = Σ(O − E)²/E = 56.25/37.5 + 56.25/22.5 + 56.25/37.5 + 56.25/22.5 = 1.5 + 2.5 + 1.5 + 2.5 = 8.0. A 2 × 2 table has (2 − 1)(2 − 1) = 1 degree of freedom, and 8.0 exceeds both 3.84 (0.05) and 6.63 (0.01).
  5. Match the method with the name attached to it. (a) Grounded theory (b) Participatory Rural Appraisal (c) Extended-case method (d) Content analysis as "objective, systematic and quantitative description of manifest content". Names: (1) Chambers (2) Berelson (3) Glaser and Strauss (4) Gluckman

    1. a-1, b-3, c-2, d-4
    2. a-3, b-4, c-1, d-2
    3. a-2, b-1, c-4, d-3
    4. a-3, b-1, c-4, d-2
    Show answer

    Answer: D — a-3, b-1, c-4, d-2

    Glaser and Strauss published The Discovery of Grounded Theory in 1967; Robert Chambers led the shift from RRA to PRA; the extended-case method belongs to Gluckman's Manchester school; Berelson's 1952 definition of content analysis is the one quoted.
  6. Which of the following are non-parametric tests? Select all that apply.

    1. Kruskal–Wallis test
    2. Mann–Whitney U test
    3. Chi-square test
    4. Student's t-test
    Show answer

    Answer: A — Kruskal–Wallis test; B — Mann–Whitney U test; C — Chi-square test

    Chi-square, Mann–Whitney and Kruskal–Wallis make no assumption of normality and work on categorical or ranked data; the t-test is parametric, assuming interval data from a normal population.
  7. A researcher wants to model whether a child is stunted (yes/no) as a function of household income, mother's schooling and caste. The appropriate technique is

    1. Simple linear regression of height on income
    2. A paired t-test between stunted and non-stunted children
    3. Multiple logistic regression, reading each coefficient as an odds ratio
    4. Pearson's correlation between stunting and caste
    Show answer

    Answer: C — Multiple logistic regression, reading each coefficient as an odds ratio

    A binary dependent variable with several predictors calls for logistic regression, which models the log-odds of the outcome; linear regression needs a continuous outcome, Pearson's r needs two interval variables, and a paired t-test needs matched measurements.
  8. 120 offspring are expected in a 3 : 1 ratio (90 : 30) but 75 : 45 are observed. What is the chi-square value for goodness of fit?

    Numerical answer — type the value.

    Show answer

    Answer: 10

    χ² = (75 − 90)²/90 + (45 − 30)²/30 = 225/90 + 225/30 = 2.5 + 7.5 = 10. On 1 degree of freedom this exceeds 3.84, so the 3 : 1 hypothesis is rejected at the 0.05 level.
  9. Assertion (A): The median, not the mean, is the appropriate average for landholding in an Indian village. Reason (R): Landholding is positively skewed, and the mean is pulled upward by a few large holdings while the median is not.

    1. Both A and R are true, and R is the correct explanation of A
    2. Both A and R are true, but R is not the correct explanation of A
    3. A is true, but R is false
    4. A is false, but R is true
    Show answer

    Answer: A — Both A and R are true, and R is the correct explanation of A

    Both are true and R explains A: in a right-skewed distribution mean > median > mode, and the median, being a positional average, is not affected by the size of the extreme values.
  10. Which statements about the techniques of analysis are correct? Select all that apply.

    1. Fairclough's critical discourse analysis works at three levels: text, discursive practice and social practice
    2. Inter-coder reliability in content analysis can be measured by Krippendorff's alpha
    3. Content analysis in Berelson's definition is a qualitative reading of latent meaning only
    4. Labov's model of oral narrative includes orientation, complicating action, evaluation and coda
    Show answer

    Answer: A — Fairclough's critical discourse analysis works at three levels: text, discursive practice and social practice; B — Inter-coder reliability in content analysis can be measured by Krippendorff's alpha; D — Labov's model of oral narrative includes orientation, complicating action, evaluation and coda

    Berelson defined content analysis as quantitative description of manifest content; the qualitative reading of latent meaning is a later extension. The statements on Labov, Fairclough and Krippendorff are correct.