What is data leakage?
Information from outside the training set influencing the model. Common causes are scaling or imputing before splitting, using features that would not exist at prediction time, and tuning hyperparameters on the test set. It produces excellent numbers and a model that fails in reality, and it is the most common serious flaw in student projects.
This comes up on Machine Learning Coursework, where it is answered in the context of the work itself.
Not what you asked?
Ask us the actual question
Send the draft, the rubric or the score report with it. You get a real answer and a price before you commit to anything.