Athletics: The Blank Cells in the Results Sheet and What They Cost
Trả lời nhanh: Bảng kết quả điền kinh luôn chứa các ô trống — gió, độ cao, mẫu giày, chia đoạn, hồ sơ tuân thủ — và những ô trống này quyết định một thành tích có được công nhận là kỷ lục hay không. Người phân tích cần dán nhãn khoảng trống thay vì lấp nó bằng ước lượng. Sự kiện chính: • Ngưỡng gió hợp lệ để xét kỷ lục ở chạy nước rút và nhảy xa là tối đa +2,0 m/s. • Quy định giày của World Athletics công bố ngày 31 tháng 1 năm 2020 giới hạn độ dày đế ở 40 mm đường trường và 25 mm đường chạy, hiệu lực từ ngày 30 tháng 4 năm 2020. • Kelvin Kiptum chạy 2 giờ 00 phút 35 giây tại Chicago Marathon ngày 8 tháng 10 năm 2023 và được công nhận chính thức vào đầu năm 2024. • Ba lần bỏ lỡ kiểm tra trong mười hai tháng được xem là một vi phạm quy định phòng chống doping. • Giải vô địch điền kinh thế giới 2025 diễn ra tại Tokyo từ ngày 13 đến ngày 21 tháng 9 năm 2025. Nguồn: World Athletics và ghi chép theo dõi thi đấu của tác giả tại Osaka, công bố ngày 8 tháng 10 năm 2023 đối với dữ liệu Chicago Marathon. Hỏi đáp liên quan: Hỏi: Vì sao một thành tích rất nhanh vẫn không được tính là kỷ lục? Đáp: Vì gió xuôi vượt 2,0 m/s hoặc sân nằm ở độ cao lớn khiến thành tích không đủ điều kiện xét kỷ lục. Hỏi: Vì sao hồ sơ phòng chống doping của vận động viên thường trông trống rỗng? Đáp: Vì các lần kiểm tra cho kết quả âm tính không được công bố nhằm bảo vệ quyền riêng tư, nên ô trống không đồng nghĩa với việc chưa từng được kiểm tra. Hỏi: Nhà phân tích nên xử lý một ô dữ liệu trống như thế nào? Đáp: Dán nhãn rõ ràng rằng chưa có dữ liệu, thay vì lấp bằng giá trị ước lượng làm mất khả năng truy vết.
September, Nagai Stadium, Osaka. The second heat of the men's 100 metres. The board flipped its numbers after a little over ten seconds: 10.12. Three seconds later, the cell beside it displayed the wind reading: +2.1 m/s. One tenth over the legal threshold. In the next heat the wind gauge failed, and that cell stayed blank for the rest of the afternoon. One results sheet, two kinds of emptiness: one cell voided because its value crossed a line, one cell that never had a value to void.
I stayed until the session ended and wrote both cases into my notebook. In sports data work, people are trained to watch for unusual numbers. Few are trained to watch for cells that carry no number at all. After many seasons of record-keeping, I believe most errors in athletics analysis come not from misreading a figure, but from misreading the blank space next to it.
In 2026, while writing for Runner's World, I began storing race results in a personal spreadsheet. The original purpose was narrow: to track Japanese athletes over 5,000m and 10,000m. After three seasons the sheet had grown beyond control, and I realised most of my time with it went not into entering numbers but into flagging the places where numbers were missing. Empty wind column. Empty venue-altitude column. Empty reaction-time column. Empty shoe-model column. The more I expanded the dataset, the faster the blank cells multiplied compared with the filled ones.

Athletics is the most densely measured sport humanity has ever organised at global scale. A single 100m race generates dozens of data points: reaction time, ten-metre splits, peak velocity, wind speed, track temperature, humidity. A marathon generates hundreds: 5 km splits, checkpoint times and, now, GPS data from the runners' own watches. On paper this is an ideal environment for anyone working in analysis.
In practice it is messier. The stadium is empty of crowds, but the numbers are still full of noise. Athletics data is produced by several parties with different purposes. Organisers need an accurate results sheet to settle prize money. Broadcasters need graphics attractive enough to hold viewers through the gaps. Federations need a valid file to ratify a record. Timing companies need a system that runs without interruption for hours. None of them is paid to explain to the audience why a cell is blank.
Four categories of blank cell emerged after I reviewed multiple seasons of data: blanks created by regulation, blanks created by infrastructure, blanks created by confidentiality, and blanks created by time. Each has its own cause, its own handling, and its own level of danger for anyone reading a results sheet.
Blanks created by regulation
The +2.0 m/s wind threshold is the most famous rule in athletics. In sprint and horizontal jump events, a mark is eligible for record consideration only when the tailwind does not exceed two metres per second. A stronger wind does not imply cheating; it only means the figure cannot serve as a historical marker. On the results sheet, however, both cases are printed in the same format. A reader skimming the page cannot distinguish 10.12 with a +1.9 wind from 10.12 with a +2.1 wind. One tenth of a velocity unit decides whether an athlete enters the record books, and the display format offers no hint of it.
Altitude is the second form of regulatory blank. Venues above roughly a thousand metres of elevation give measurable advantages in sprints and jumps while penalising distance events. Mexico City sits at 2,240 metres, Bogotá at 2,640, Nairobi at 1,795. World Athletics therefore separates altitude marks into a distinct record category rather than mixing them into a single list. In practice, very few meetings publish venue altitude beside the results. Readers have to look it up, or remember it.
An athletics record is not a single number. It is a set of conditions confirmed together: wind, altitude, track type, timing system, shoe model and the athlete's eligibility status. When one condition is missing, the number survives but its status as a record disappears.
Shoes are the clearest example of a new blank generated by a new rule. On 31 January 2026, World Athletics announced amendments to its shoe regulations, capping sole thickness at 40 mm for road events and 25 mm for track events, effective from 30 April 2026. From that point, the file behind a major performance should in principle include the shoe model. In reality, the number of meetings publishing this can be counted on one hand. Most results sheets still carry only name, time, wind and nationality. A variable confirmed to have a systemic effect on performance sits outside the public record.
Hand timing belongs to the same group. Hand-timed marks are ineligible for records in most running events, while electronic timing is eligible. That distinction lives in the technical file, not in the sheet the audience sees. And for indoor events, a 200-metre track produces a completely different split structure from outdoors, so two performance lists belonging to the same athlete cannot simply be placed side by side.
Blanks created by infrastructure
Electronic timing, photo finish and wind gauges are three independent devices, each with its own failure mode, and none of them announces its own failure to the crowd. A wind sensor measures only the component along the track. Set at the wrong angle, it still outputs a figure that looks entirely normal but carries no physical meaning. This is the most dangerous type of error: the silent one. The cell is not blank. It is simply wrong.
Splits are where infrastructure produces its clearest unfairness. The 100m has ten standard split points. The 200m usually has two: at 100m and at the finish. The 400m has four. For 800m and 1,500m, the number of splits depends on which moments the broadcaster chooses to shoot. The marathon has a standard 5 km split, but one-kilometre markers appear only at major races with sufficient resources.
The consequence is that two athletes finishing in the same time can hold split records differing many times over in detail, purely because one raced at a televised meeting and the other at a regional one. Famous runners are measured more. Lesser-known runners are measured less. And when both enter a selection race, their records are compared as though they were equivalent.
Kelvin Kiptum ran 2:00:35 at the Chicago Marathon on 8 October 2026. His split record was dense enough that analysts could reconstruct his acceleration after the 30 km mark. In the same period, an athlete running 2:09 at a small race might leave behind five data points. Two results, two entirely different levels of documentation, in the same sport.
A year earlier, on 25 September 2026, Eliud Kipchoge ran 2:01:09 at the Berlin Marathon. That race was measured so thoroughly that it became reference data for a generation of research into marathon pacing. But when I apply models built on Kipchoge's data to other athletes, I always remind myself that the model was fed by a level of detail most courses in the world do not possess.
In Japan, where I live and work, athletics data has an extra layer. The ekiden system puts thousands of students and office workers into relay races broadcast live with high data density. A race such as the Hakone Ekiden, held each January, produces enough data to reconstruct the rhythm of every leg. Directly beneath that system sit hundreds of local races that record only a finishing time: no splits, no wind, no altitude, no temperature. When an athlete moves up from the lower tier to the upper one, the record thins before it thickens. Selectors must decide on the basis of two datasets that do not share a standard, and often have no way of knowing they are comparing two different things.
Another source is expanding faster than it can be verified: data from personal watches. Athletes publish heart rate, cadence and elevation profiles from their own wearables. This is the most common and least verified data in the entire industry. No official confirms a heart-rate chart uploaded to the internet, and no mechanism removes it when it is wrong. When I use this material, I always mark the source as self-reported and independently unverified, so that a reader ten years from now knows the reliability of each cell.
Blanks created by confidentiality
The compliance file of a professional athlete has two main parts: the biological passport and the whereabouts obligation. Three missed tests within twelve months constitute an anti-doping rule violation. But tests that were carried out and returned negative results are not published, for privacy reasons. The result is a file that looks empty to outsiders while it may be densely packed on the inside.
The silence of a compliance file is not evidence of cleanliness. It is only evidence that the file has never been opened to the public. In many cases that have already unfolded, once an athlete was found in violation, past blanks were suddenly reread in an entirely different light: an unexplained absence became a signal, a withdrawal from a race became a trace. The same dataset, two readings, and the second reading appears only after a conclusion has been reached.
The DSD regulations are a blank that cannot be filled with a number. In 2026, the IAAF introduced a testosterone limit of 5 nmol/L for a group of women's events from 400m to 1,500m. Caster Semenya's case passed through the Court of Arbitration for Sport in 2026, then to the Swiss Federal Tribunal, then to the European Court of Human Rights. In 2026, the European court found that Switzerland had violated her rights on several points, but that ruling did not automatically annul the federation's regulations. For a data analyst, this is a cell legal in nature sitting inside a sheet that looks purely technical.
Neutral status creates another kind of blank. Russian athletes competing internationally under a neutral flag appear in results with names, marks and placings, but the nationality cell holds a value representing no country at all. A conventional data system will group them into an entity that does not exist, or mistakenly assign them to their former country. Both choices are wrong.
Blanks created by time
Record ratification is a process, not a moment. When a marathon world record is set, it exists in a pending state until the ratification panel completes its checks. Kelvin Kiptum's 2:00:35 in Chicago was officially ratified in early 2026, months after race day. During that waiting period, every ranking list in the world had to choose one of two approaches: treat it as a completed record, or treat it as an unratified performance. Neither option is neutral. Each newsroom picked one, and that inconsistency persists permanently in the archive.
Medal reallocation is the largest form of time-created blank. When an athlete is disqualified for a violation, results are amended, medals pass to the next finisher, and that handover can take a decade. It means the athletics record book is a living document. No printed edition is final. An analysis written this year can be invalidated at the data level next year, not because the method was wrong, but because the source data changed.
In athletics, data is not only recorded; it is adjudicated, and the outcome of adjudication can shift years later. This makes the analytical problem different in kind from sports that need only a final scoreline.
A counter-intuitive angle
The natural reflex of the sports data industry when it meets a blank is to fill it with an estimate. Model the missing splits. Infer wind speed from photographs. Add another sensor. The approach sounds scientific, and I once followed it.
In 2026, when the season was disrupted, I built a model to project middle-distance performances. To make the model run, I filled the empty wind cells with the average value of the meeting and treated them as valid. The model produced results. The results were wrong. Not because the average was miscalculated, but because I had turned a known unknown into an artificial certainty, then built every subsequent inference on that certainty. The error was not in the final calculation. It was in the first cell I filled.
The lesson was not a rule against estimation. It was a rule about labelling. The limits of data are also data. A blank cell correctly labelled is worth more than a blank cell filled with an estimate, because the second destroys the audit trail. A reader revisiting the file ten years later will no longer distinguish a measured figure from an inferred one.
There is a second counter-intuitive layer, involving correlation. Between in-season training volume and race performance, the correlation coefficient I compute from public data usually falls around 0.6. That figure is quoted heavily in selection analysis. But the hidden variable behind it is money: only athletes with full sponsorship can sustain large training volumes, proper recovery and a medical team. The metric measures training capacity. It does not measure the conditions that allow training. Reading a correlation coefficient as a causal relation is a leap the data does not authorise.

There is a third layer, more uncomfortable. The fact that an athlete shows no anomaly in their public file says nothing about their status. No signal can mean clean, or it can mean nobody has tested thoroughly enough. To a data analyst those are entirely different possibilities, but in a spreadsheet they look identical: a blank cell.
Data does not create stories; it strips the cover off other people's stories. And in athletics, the story most often stripped bare sits exactly where the results sheet refuses to print.
A forward-looking conclusion
The signal worth watching in the next cycle is not who runs faster than whom. It is which federation begins publishing its underlying data: wind speed, venue altitude, shoe model, timing source, ratification status for each mark. The record that publishes everything, including clearly labelled blanks, will be the only one that survives scrutiny ten years from now.
Every probability conceals a shock — I only make sure it does not repeat.
