Đề IELTS Reading · IELTS 8020

What Tests Can and Cannot Measure

← Tất cả đề Reading
IELTS Academic Reading Band 6.5-8.0 13 câu Bài đọc ~898 từ 14 phút Đề 8020 tự biên soạn

Trang này có toàn văn bài đọcđủ 13 câu hỏi đúng như trong phòng thi, chia theo dạng: True/False/Not Given · Điền từ · Multiple choice. Đáp án và lời giải từng câu không in ở đây — bạn làm bài trên máy rồi hệ thống chấm ngay khi nộp và giải thích vì sao mỗi câu đúng hoặc sai. Làm trước, đọc lời giải sau thì mới biết mình sai ở đâu; đọc đáp án trước thì đề coi như hỏng.

Làm đề này trên máy, chấm ngay khi nộp

Đúng định dạng thi máy, có đồng hồ. Nộp xong hiện đáp án kèm lời giải từng câu. Không cần trả phí.

Vào làm đề này →

Bài đọc

AThe case for examining every candidate on the same paper under the same conditions is easy to state: the resulting numbers can be compared without reference to the reputation of the school that produced them. It is a promise of fairness that has proved remarkably durable, and since the first large-scale administration of a common university entrance paper in 1926 the instrument has spread to almost every education system in the world. What is less often noticed is that a test measures nothing directly. It elicits a small sample of behaviour, from which an inference is drawn about a capacity nobody can observe, and it is the inference, rather than the paper itself, that may be sound or unsound. The distinction sounds pedantic and turns out to be the source of most of the argument that follows. A paper yielding consistent marks from one occasion to the next may still support a conclusion that is wholly unwarranted, in much the way that a reliable thermometer tells us nothing whatever about the depth of a river.

BThe first serious challenge came from the observation that marks can rise without the capacity behind them rising as well. Between 2011 and 2016 the Wrenford authority recorded an increase of 14 points in the average result on its state examination, a rise reported at the time as evidence that its schools had improved. Ines Baruch, of the Halvard Institute of Measurement, arranged for a sample of the same pupils to sit a low-stakes audit paper covering the identical syllabus, on which the average moved by 1.4 points across the same five years. She argues that the divergence measures neither learning nor cheating but familiarity: what improves is performance on the particular questions a system has learnt to expect. Her term for the phenomenon is score inflation, and she insists that it is invisible from inside a single testing regime, since nothing within that regime supplies an independent yardstick. The audit paper supplied one, and the two lines parted. Wrenford's schools were never accused of dishonesty; the drilling that produced the gain was, in the main, ordinary teaching aimed at a predictable target.

CDefenders of the instrument reply that its critics have been misreading the statistics rather than exposing a fault in the papers. Tomás Rautio points out that the correlations usually quoted between entrance marks and later results are calculated only among candidates who were admitted, a group from which the weakest performers have already been removed. Restricting the range in this way shrinks any correlation mechanically, whatever the underlying relationship may happen to be. When he applied the standard statistical correction to admissions data from 2014, the coefficient rose from 0.34 to 0.51, which he maintains is respectable for a single predictor of something as loosely defined as academic success. Rautio concedes, however, that the correction rests on assumptions about the applicants who were turned away and whose later performance nobody can observe, and that his own figures come from one institution over three years. Assumptions of that kind cannot be tested directly. He therefore presents the corrected number as a ceiling rather than as a measurement, a distinction his more enthusiastic supporters have not always preserved.

DMore consequential than what a test measures is what it causes. Once results are used to judge institutions rather than to place individuals, the incentive is no longer to teach the subject but to teach the paper, and the two overlap only in part. Teachers surveyed in the Alcombe district reported giving an average of nine weeks a year to examination technique, and their pupils' marks rose by 11 per cent while performance on tasks demanding extended writing did not move at all. Baruch regards this as the clearest demonstration that a rise in marks and a rise in competence are separable quantities. A related difficulty concerns individual questions rather than whole papers: analysis at the level of the item has repeatedly found that a small number of questions behave differently for pupils whose first language is not the language of the paper, but only where the question depends on a context outside the syllabus, such as a description of an unfamiliar sport. Removing those items is straightforward once they have been identified, and most examining bodies now do so.

EProposals for reform therefore tend to modify the instrument rather than abandon it. Adaptive testing, in which each question is selected according to how the candidate answered the last, reaches a comparable precision in about half the time, though it demands banks of thousands of calibrated items and sits awkwardly with the principle that everyone should visibly face the same paper. Comparative judgement, in which examiners rank scripts against one another instead of scoring them against a scheme, produces markedly steadier orderings for extended writing, at some cost in transparency. Neither reform touches the difficulty Baruch identified, since both remain vulnerable to preparation aimed at the format. The most defensible position is probably that the fault lies less in the instrument than in the weight placed upon it. A paper used to allocate places among candidates of similar preparation is doing what it was designed to do; the same paper used to rank schools, reward teachers and direct budgets is being asked to carry a load no small sample of behaviour can bear. That load, and not the paper, is what reform should lighten.

Câu hỏi (13 câu)

Questions 1–5 · TRUE / FALSE / NOT GIVEN

Do the following statements agree with the information given in the passage? Write TRUE if the statement agrees with the information, FALSE if the statement contradicts the information, NOT GIVEN if there is no information on this.

  1. 1.Questions of every sort have been shown to work differently for candidates whose mother tongue differs from the one an examination is written in.
  2. 2.Baruch reads the Alcombe findings as a further case of the split between marks and ability that she detected in a different authority.
  3. 3.Rautio describes the adjusted figure of 0.51 as an accurate estimate of how well entrance marks predict later work.
  4. 4.The statistical adjustment used by Rautio has since been taken up by the organisations that build tests of the adaptive kind.
  5. 5.The changes to examining now under discussion would leave untouched the problem created by preparation directed at the format.

Questions 6–9 · Sentence completion

Complete the sentences below. Choose NO MORE THAN TWO WORDS from the passage for each answer.

  1. 6.Baruch's name for a climb in marks unaccompanied by any growth in ability is ________.
  2. 7.The gap only became visible because the same pupils also sat an ________ that carried no consequences for them.
  3. 8.Staff in one district reported that a large part of each year was given over to ________.
  4. 9.Placing scripts in order against each other, a procedure called ________, gives steadier outcomes for longer written answers.

Questions 10–13 · Multiple choice

Choose the correct letter, A, B, C or D.

  1. 10.Why does the writer mention a thermometer in the first paragraph?
    1. A. To suggest that examiners ought to record conditions as carefully as scientists.
    2. B. To show that steady readings can sit alongside a conclusion worth nothing.
    3. C. To argue that measurement in the physical sciences is far more exact.
    4. D. To illustrate how seldom instruments of any sort are checked for accuracy.
  2. 11.Rautio's argument about the coefficient of 0.34 depends on the claim that
    1. A. those left out of the calculation would have done unusually well afterwards.
    2. B. the correction he used was devised for a different sort of examination.
    3. C. entrance papers foresee later results better than any rival single measure.
    4. D. the group behind the calculation is not the group that originally applied.
  3. 12.What does the writer indicate about getting rid of questions that behave unevenly?
    1. A. It lies beyond the resources of most bodies that set examinations.
    2. B. It matters far less than the wider question of how results are used.
    3. C. It causes little trouble once such questions have been picked out.
    4. D. It is needed only where candidates share a single first language.
  4. 13.Which statement best sums up the writer's own position at the end of the passage?
    1. A. Examinations ought to be replaced wherever a workable alternative exists.
    2. B. The instrument is sound enough, but the uses made of it often are not.
    3. C. Consistency has been overvalued and soundness left almost entirely aside.
    4. D. Inflated marks make any comparison between schools impossible in practice.
Tự chấm giờ: đề này gợi ý 14 phút. Trong bài thi Reading thật bạn có 60 phút cho 3 passage và 40 câu, nên hãy tập bám sát mốc thời gian ngay từ khi luyện — hết giờ là kiểu mất điểm phổ biến nhất của phần Reading.

Cách làm các dạng câu có trong đề này

TRUE / FALSE / NOT GIVEN

FALSE nghĩa là bài nói NGƯỢC LẠI, không phải bài không nói. Còn NOT GIVEN nghĩa là bài im lặng về chuyện đó. Quy tắc tự kiểm rẻ nhất: khi định trả lời FALSE, hãy chỉ tay vào đúng cụm từ trong bài mâu thuẫn với phát biểu — không chỉ ra được thì đáp án là NOT GIVEN.

Các câu theo đúng thứ tự xuất hiện trong bài đọc, nên khi đã định vị được câu 3 và câu 5 thì câu 4 chắc chắn nằm giữa hai chỗ đó. Đừng đọc lại cả bài cho từng câu.

Đọc kỹ hơn: phân biệt True/False/Not Given với Yes/No/Not Given.

Điền từ (Sentence / Summary / Note completion)

Đọc giới hạn số từ trong câu lệnh trước khi làm câu đầu tiên. Viết quá giới hạn là sai, kể cả khi nội dung đúng. Từ ghép có gạch nối tính là một từ; mạo từ a, the vẫn tính là một từ nên bỏ được thì nên bỏ.

Trước khi đi tìm, hãy đoán từ loại cho mỗi chỗ trống dựa vào ngữ pháp của câu: danh từ, số, hay động từ. Việc này biến bài đọc từ "đọc xem có gì" thành "đọc để xác nhận cái mình đang chờ". Chính tả và số ít số nhiều đều bị chấm.

Đọc kỹ hơn: luật số từ và bẫy điền từ.

Multiple Choice

Loại hai đáp án sai trước, rồi mới so hai đáp án còn lại — đừng cố tìm đáp án đúng ngay từ đầu. Đáp án sai của IELTS thường sai vì một chữ: một trạng từ tuyệt đối (always, only), một chủ thể bị đổi, hoặc một quan hệ nhân quả bài không hề khẳng định.

Đáp án đúng gần như luôn là bản diễn đạt lại của câu trong bài, không phải bản chép nguyên chữ. Phương án dùng lại nhiều từ y hệt bài đọc thường là bẫy.

Đọc kỹ hơn: các dạng câu hỏi Reading khác.

Sẵn sàng làm thử?

Làm xong sẽ thấy đáp án, lời giải từng câu và chỗ trong bài đọc quyết định đáp án đó.

Làm đề "What Tests Can and Cannot Measure" →

Đề Reading khác cùng mức

Xem toàn bộ kho đề IELTS Reading, hoặc vào kho đề luyện tập để lọc theo kỹ năng và dạng câu. Đang cần một khung học tổng thể thì xem lộ trình tự học IELTS.