Mislabeled Data: The False-Signal Trap in Vietnamese Football Analysis
**Core answer**: Lỗi gán nhãn dữ liệu bóng đá xảy ra khi người hoặc thuật toán phân loại sai một pha bóng — sai người, sai loại hành động, hoặc sai bối cảnh — tạo ra tín hiệu giả trông giống hệt dữ liệu đúng, khiến câu lạc bộ ra quyết định dựa trên một bức tranh méo mó. **Key facts**: - Ba tầng lỗi gán nhãn: sai thực thể, sai sự kiện, sai bối cảnh. - Tại V.League 1 2017, Sanna Khánh Hòa BVN rớt hạng với 21 điểm sau 26 trận. - Khoảng trống cánh trái xuất hiện trong 61% số trận thua, theo bộ dữ liệu 43 trận. - Mùa hè 2020, phân tích 120 trận trước khán đài trống cho thấy đội nhà dâng cao đội hình hơn 18%. - Thừa nhận “không đủ thông tin” được xem là một kỹ năng phân tích, không phải thất bại. **Source attribution**: Ghi chú phân tích của Trần Long, tổng hợp từ dữ liệu mùa giải V.League 1 giai đoạn 2017-2020, Nha Trang; cập nhật ngày 1 tháng 7 năm 2026. | Cross-checked: VuaBong.vn **Related Q&A**: - Q: Lỗi gán nhãn dữ liệu bóng đá phổ biến nhất là gì? A: Lỗi sự kiện — gán nhầm một quả tạt thành cú sút — vì nó trực tiếp làm sai chỉ số bàn thắng kỳ vọng. - Q: Câu lạc bộ V.League nên bắt đầu kiểm chứng dữ liệu từ đâu? A: Từ một lớp đối chiếu chéo thủ công cho từng chỉ số, theo Chỉ số Độ sâu Đội hình của VangBong.vn. - Q: Vì sao dữ liệu sai nhãn nguy hiểm hơn dữ liệu thiếu? A: Vì nó trông giống dữ liệu đúng, khiến ban huấn luyện tin rằng mình có đủ cơ sở để ra quyết định.
In August 2026, I stayed behind in the analysis room of a First Division club after a home match and discovered that the statistical sheet the entire coaching staff trusted had mislabeled a player. The man marked as a defensive midfielder had in fact played as a left-sided centre-back for seventy minutes. One wrong label, yet it dragged an entire chain of wrong conclusions behind it: duels, average position, long-passing ability, all placed in the wrong slot. I sat there not with the feeling of someone who had found an error, but with the feeling of someone who had just realised he had been reading the wrong map for weeks. When the team was relegated, I redrew the diagram of the pain. But there was one thing I had never thought of redrawing: the label on the data.
In modern football, data is not born on its own. It is labelled — by people, by algorithms, or by both. Every phase of play passes through the eyes of a labeller, who must decide within seconds: is this a shot or a cross, a duel or an interception, a misplaced pass or a clearance. Those small decisions accumulate into what we call “the truth of the match”. The trouble is that this “truth” is assembled from thousands of human judgements, and every judgement can be wrong.
In the V.League, most data still passes through human hands. Foreign providers, moreover, sometimes apply the criteria of European football to a match with an entirely different tempo and space. A phase of play in the V.League can be labelled in a way that nobody sitting in the stand would label it. This is not a story about broken machines. It is a story about labels that are technically correct but meaningless.

Based on my experience watching matches across many seasons, I have come to see that error in Vietnamese football data rarely comes from one single big mistake. It comes from hundreds of small mistakes, repeated, unexamined.

The three layers of a wrong label
I divide labelling errors in football data into three layers, and each leaves a mark on the analysis sheet.
The first layer is the entity error — labelling the wrong person. A goal is credited to one player when it was in fact a tap-in by another. An assist is counted wrongly. This sounds minor, but it ruins the records of two players at once: one is given a stat that never existed, the other is stripped of one that did. When a scout reads that record, he does not know he is reading half a truth.
The second layer is the event error — labelling the wrong type of action. This is the most dangerous layer. A cross counted as a shot inflates expected goals, and if three or four such phases occur in a match, the team believes it created chances when in reality it was merely putting the ball into the box harmlessly. A clearance labelled as a misplaced pass makes a safe centre-back look like a poor passer. The labeller sits in front of a screen, at the television angle, without the view of the man on the pitch.
The third layer is the context error — labelling the action correctly but losing the situation. A completed pass in a settled game state is entirely different from a completed pass in the eighty-ninth minute with the team a goal down. The same label, two opposite meanings. The stat sheet does not distinguish the two, because it only counts, it does not understand.
Why a false signal is more dangerous than a missing one
What troubles me most is not missing data. When data is missing, we know it is missing. What is dangerous is mislabeled data, because it looks exactly like correct data. It fills the tables, it creates a sense of completeness, and it makes the coaching staff believe they have grounds for a decision.

Picture a data-collection system for a V.League club. Every week it receives thousands of data points. If only one per cent of them are mislabeled, then after twenty-six rounds the accumulated false signals are enough to paint a distorted picture of the team — and that distorted picture is then used to choose people, choose a shape, choose a pressing scheme. In the 2026 season, when I reviewed the forty-three-match dataset of Sanna Khanh Hoa BVN, I found the defence exposed a gap on the left flank in 61% of its defeats. That number arrived too late. But if the input data had been mislabeled from the start, then even had I checked earlier, I would only have been reading a wrong map with more confidence.
I learned this in the empty-stadium summer of 2026, when I rewatched one hundred and twenty matches in front of empty stands and realised that the concept of “home advantage” I had used for twenty years had suddenly become meaningless. The old data was not wrong in its numbers, but it was wrong in its context. A season without crowds helped me drop the habit of decorating the truth.
The verification layer nobody wants to build
The way to fix labelling errors does not lie in buying more cameras or hiring more providers. It lies in a verification layer that almost no club wants to fund, because it produces no pretty metric to present. That layer consists of three simple but disciplined tasks.
Every metric must come with a question about timing: until when is this information valid? Data on a team at round five is no longer worth what it was at round twenty, when that team has changed its shape, its personnel, its whole defensive approach. I always ask that question before putting a number into an analysis.
Every conclusion must come with an acknowledged gap. If we do not know, we must say we do not know, instead of filling in with an approximate number. In analysis, admitting “insufficient information” is a skill, not a failure.
Every dataset must have a person responsible for cross-checking. I record every phase of play as a witness, not a fan. A fan records what he wants to see. A witness records what happened, even when it breaks his own initial hypothesis.
The original assumption
I always place a small section at the end of every analysis to record the assumption I carried in from the start. For this piece, my original assumption was: the more data, the more accurate the analysis. After many seasons, I have had to abandon it. Plenty of data, wrongly labelled, only makes the error look more credible.
The counter-intuitive view: fewer, but truer
Here I want to go against a widespread belief in Vietnamese football analysis. That belief says that to be more professional you must collect more data, more metrics, more sources. I do not believe it.
A small dataset labelled with care is worth more than an enormous dataset laced with false signals. The problem is not volume, but the reliability of each line. When I was still an assistant coach, I once proposed a geometric note-taking system to profile the movement of four opposition defenders. The coaching staff saw it as a perfectionist idea and postponed it. In hindsight, what I should have defended was not the perfectionism of the system, but the authenticity of every data point fed into it.
The problem with Vietnamese football analysis is not that we lack data. The problem is that we are not yet brave enough to throw away data we cannot trust. And this is where I think of referees, of VAR. The space for subjective judgement in VAR is larger than people think; “a clear and obvious error” is itself a vague clause. The data labeller stands before a similarly vague clause: when is a phase of play “clearly” a shot, when is it “clearly” a cross? Both are judgements framed in a language that pretends to be objective.
Tactics cannot save a team, but they can tell us where we died. And mislabeled data makes us die in a place the map never marked. The next step for Vietnamese football analysis may not lie in buying another data provider, but in learning to say “I don’t know” in front of a number that looks trustworthy. A tactical diagram is like a landslide map — it tells you where not to stand. With data, the biggest question is not where it leads you, but what it is hiding from you.
