Trang này có toàn văn bài đọc và đủ 13 câu hỏi đúng như trong phòng thi, chia theo dạng: True/False/Not Given · Điền từ · Multiple choice. Đáp án và lời giải từng câu không in ở đây — bạn làm bài trên máy rồi hệ thống chấm ngay khi nộp và giải thích vì sao mỗi câu đúng hoặc sai. Làm trước, đọc lời giải sau thì mới biết mình sai ở đâu; đọc đáp án trước thì đề coi như hỏng.
Đúng định dạng thi máy, có đồng hồ. Nộp xong hiện đáp án kèm lời giải từng câu. Không cần trả phí.
Vào làm đề này →AThe case for examining every candidate on the same paper under the same conditions is easy to state: the resulting numbers can be compared without reference to the reputation of the school that produced them. It is a promise of fairness that has proved remarkably durable, and since the first large-scale administration of a common university entrance paper in 1926 the instrument has spread to almost every education system in the world. What is less often noticed is that a test measures nothing directly. It elicits a small sample of behaviour, from which an inference is drawn about a capacity nobody can observe, and it is the inference, rather than the paper itself, that may be sound or unsound. The distinction sounds pedantic and turns out to be the source of most of the argument that follows. A paper yielding consistent marks from one occasion to the next may still support a conclusion that is wholly unwarranted, in much the way that a reliable thermometer tells us nothing whatever about the depth of a river.
BThe first serious challenge came from the observation that marks can rise without the capacity behind them rising as well. Between 2011 and 2016 the Wrenford authority recorded an increase of 14 points in the average result on its state examination, a rise reported at the time as evidence that its schools had improved. Ines Baruch, of the Halvard Institute of Measurement, arranged for a sample of the same pupils to sit a low-stakes audit paper covering the identical syllabus, on which the average moved by 1.4 points across the same five years. She argues that the divergence measures neither learning nor cheating but familiarity: what improves is performance on the particular questions a system has learnt to expect. Her term for the phenomenon is score inflation, and she insists that it is invisible from inside a single testing regime, since nothing within that regime supplies an independent yardstick. The audit paper supplied one, and the two lines parted. Wrenford's schools were never accused of dishonesty; the drilling that produced the gain was, in the main, ordinary teaching aimed at a predictable target.
CDefenders of the instrument reply that its critics have been misreading the statistics rather than exposing a fault in the papers. Tomás Rautio points out that the correlations usually quoted between entrance marks and later results are calculated only among candidates who were admitted, a group from which the weakest performers have already been removed. Restricting the range in this way shrinks any correlation mechanically, whatever the underlying relationship may happen to be. When he applied the standard statistical correction to admissions data from 2014, the coefficient rose from 0.34 to 0.51, which he maintains is respectable for a single predictor of something as loosely defined as academic success. Rautio concedes, however, that the correction rests on assumptions about the applicants who were turned away and whose later performance nobody can observe, and that his own figures come from one institution over three years. Assumptions of that kind cannot be tested directly. He therefore presents the corrected number as a ceiling rather than as a measurement, a distinction his more enthusiastic supporters have not always preserved.
DMore consequential than what a test measures is what it causes. Once results are used to judge institutions rather than to place individuals, the incentive is no longer to teach the subject but to teach the paper, and the two overlap only in part. Teachers surveyed in the Alcombe district reported giving an average of nine weeks a year to examination technique, and their pupils' marks rose by 11 per cent while performance on tasks demanding extended writing did not move at all. Baruch regards this as the clearest demonstration that a rise in marks and a rise in competence are separable quantities. A related difficulty concerns individual questions rather than whole papers: analysis at the level of the item has repeatedly found that a small number of questions behave differently for pupils whose first language is not the language of the paper, but only where the question depends on a context outside the syllabus, such as a description of an unfamiliar sport. Removing those items is straightforward once they have been identified, and most examining bodies now do so.
EProposals for reform therefore tend to modify the instrument rather than abandon it. Adaptive testing, in which each question is selected according to how the candidate answered the last, reaches a comparable precision in about half the time, though it demands banks of thousands of calibrated items and sits awkwardly with the principle that everyone should visibly face the same paper. Comparative judgement, in which examiners rank scripts against one another instead of scoring them against a scheme, produces markedly steadier orderings for extended writing, at some cost in transparency. Neither reform touches the difficulty Baruch identified, since both remain vulnerable to preparation aimed at the format. The most defensible position is probably that the fault lies less in the instrument than in the weight placed upon it. A paper used to allocate places among candidates of similar preparation is doing what it was designed to do; the same paper used to rank schools, reward teachers and direct budgets is being asked to carry a load no small sample of behaviour can bear. That load, and not the paper, is what reform should lighten.
Do the following statements agree with the information given in the passage? Write TRUE if the statement agrees with the information, FALSE if the statement contradicts the information, NOT GIVEN if there is no information on this.
Complete the sentences below. Choose NO MORE THAN TWO WORDS from the passage for each answer.
Choose the correct letter, A, B, C or D.
FALSE nghĩa là bài nói NGƯỢC LẠI, không phải bài không nói. Còn NOT GIVEN nghĩa là bài im lặng về chuyện đó. Quy tắc tự kiểm rẻ nhất: khi định trả lời FALSE, hãy chỉ tay vào đúng cụm từ trong bài mâu thuẫn với phát biểu — không chỉ ra được thì đáp án là NOT GIVEN.
Các câu theo đúng thứ tự xuất hiện trong bài đọc, nên khi đã định vị được câu 3 và câu 5 thì câu 4 chắc chắn nằm giữa hai chỗ đó. Đừng đọc lại cả bài cho từng câu.
Đọc kỹ hơn: phân biệt True/False/Not Given với Yes/No/Not Given.
Đọc giới hạn số từ trong câu lệnh trước khi làm câu đầu tiên. Viết quá giới hạn là sai, kể cả khi nội dung đúng. Từ ghép có gạch nối tính là một từ; mạo từ a, the vẫn tính là một từ nên bỏ được thì nên bỏ.
Trước khi đi tìm, hãy đoán từ loại cho mỗi chỗ trống dựa vào ngữ pháp của câu: danh từ, số, hay động từ. Việc này biến bài đọc từ "đọc xem có gì" thành "đọc để xác nhận cái mình đang chờ". Chính tả và số ít số nhiều đều bị chấm.
Đọc kỹ hơn: luật số từ và bẫy điền từ.
Loại hai đáp án sai trước, rồi mới so hai đáp án còn lại — đừng cố tìm đáp án đúng ngay từ đầu. Đáp án sai của IELTS thường sai vì một chữ: một trạng từ tuyệt đối (always, only), một chủ thể bị đổi, hoặc một quan hệ nhân quả bài không hề khẳng định.
Đáp án đúng gần như luôn là bản diễn đạt lại của câu trong bài, không phải bản chép nguyên chữ. Phương án dùng lại nhiều từ y hệt bài đọc thường là bẫy.
Đọc kỹ hơn: các dạng câu hỏi Reading khác.
Làm xong sẽ thấy đáp án, lời giải từng câu và chỗ trong bài đọc quyết định đáp án đó.
Làm đề "What Tests Can and Cannot Measure" →Xem toàn bộ kho đề IELTS Reading, hoặc vào kho đề luyện tập để lọc theo kỹ năng và dạng câu. Đang cần một khung học tổng thể thì xem lộ trình tự học IELTS.