AI coding agents can already do many tasks in a data science project.
They can explore the dataset, prepare the features, train several models, and save the final pipeline. With a simple prompt, we can build a complete machine learning workflow quickly.
However, building a model is different from understanding the problem behind it.
I experienced this difference when I asked an AI agent to build a customer churn model. I ran the experiment twice with the same dataset, environment, and prompt. The only difference was the project context available to the agent.
Both agents created a working pipeline. Both avoided data leakage, used time-based validation, and selected logistic regression. However, they did not evaluate the same problem.
This is where context engineering could become useful.
In this article, I want to explore how we can give an AI agent the data-science context it needs. I also performed a small experiment to see what happened when the project included a prediction contract.
Let’s get into it.
AI Agents in the Data Science Workflow
AI agents have become more capable of working independently inside a project. They are no longer limited to suggesting a single Python function. The agent could inspect files, write code, run the training process, fix errors, and save the model artifacts.
Recent research also evaluates this capability in a data-science setting. DataSciBench evaluates language models using realistic data-science tasks rather than only isolated code generation.
Another benchmark, AgentDS, compares AI-only and human-AI workflows across domain-specific data-science challenges. The agents produced executable pipelines and ran iterative experiments. However, they often returned to generic modeling strategies when the task required domain knowledge.
A similar situation appeared in SciAgentArena. The agent became more useful when the analysis workflow and evaluation criteria were clearly specified. The performance was less reliable when the research question was open-ended.
For me, the results show that the coding capability is only one part of the workflow.
The AI agent might know how to build a churn model, but it does not automatically understand what churn means for our project.
The Data Science Problem Is More Than the Dataset
Let’s use customer churn prediction as an example.
The dataset might contain a snapshot_date, customer information, and a target named churned_next_30d. From these columns, the AI agent could identify the prediction target, encode categorical variables, and notice that the target was imbalanced.
However, several important questions were still unanswered.
When would the prediction happen? Would the customer be scored at the start or the end of the month? Was cancellation_date already available during the prediction? Should the model use random validation or future data? How many customers could the retention team contact?
Another question: Was the model predicting churn risk, or estimating the effect of sending a retention offer?
These questions would change how we prepare and evaluate the model. However, we could not answer them reliably from the column names.
To provide the missing information, I created a prediction contract for the project.
# Customer churn prediction contract
At the end of every calendar month, the retention team scores customers
with an active subscription.
The target indicates whether the customer cancels during the 30 days
following snapshot_date.
Only information known by the end of snapshot_date may be used.
Cancellation dates, the account status after 30 days, and retention-offer
outcomes are not available at prediction time.
Use records before July 2025 for development. Records from July through
December 2025 are the untouched final evaluation period.
Average precision is the primary metric. The retention team can contact
only 10% of customers, so also report recall among the 10% with the
highest predicted risk.
The evaluation measures risk ranking. It does not estimate the causal
effect of contacting a customer.The document did not teach the agent how to create a Scikit-Learn pipeline. The agent already knew how.
Instead, the prediction contract explained decisions the agent should not make on its own.
Connecting the Project Context with the AI Agent
Repository-level instructions are one way to connect the project knowledge with an AI agent.
VS Code supports several instruction files depending on the selected agent, including .github/copilot-instructions.md, AGENTS.md, CLAUDE.md, and scoped instruction files. The VS Code documentation describes this as a context-engineering workflow, where project knowledge stays available across agent interactions.
For the experiment, I used a short AGENTS.md file.
# Project instructions
Before changing data preparation, feature engineering, training, or
evaluation, read docs/prediction-contract.md.
Treat the prediction moment, target window, unavailable information,
validation design, and evaluation metrics in that document as binding
project constraints.
Do not silently replace them with generic machine-learning defaults.
The final implementation must be runnable from the command line. Save the
fitted pipeline and a JSON metrics file, then run the finished command.The AGENTS.md file was intentionally short. Its purpose was to tell the agent where the important context was and how to use it.
The detailed information stayed inside docs/prediction-contract.md. OpenAI also described a similar pattern in its harness-engineering article, where a short AGENTS.md file points the agent toward deeper project knowledge.
For a data-science project, the deeper knowledge could include the prediction definition, data availability, validation design, metric reasoning, and operational constraints.
The flow was simple. The AI agent read AGENTS.md, while AGENTS.md directed it to the prediction contract. The contract then provided the data-science decisions used during model development.
Experiment Setup
For the experiment, I generated a synthetic customer-month dataset with 9,000 rows. The data covered January 2024 to December 2025, and the overall churn rate was around 2.38%.
The dataset also contained several risky columns:
cancellation_date
account_status_30d
retention_offer_accepted
The first two columns perfectly revealed the target because they were created after the prediction moment. Meanwhile, retention_offer_accepted represented an action that happened after customer scoring.
I then prepared two separate project repositories. Both had the same dataset, package versions, README file, and prompt. However, only the second repository contained AGENTS.md and docs/prediction-contract.md.
Both experiments used OpenAI Codex CLI 0.147.0 with the gpt-5.6-sol model and medium reasoning effort. Each condition was run once in a fresh session.
The prompt was the following:
Build a complete, reproducible customer churn model using the data in this
repository. Train and evaluate the model, save the fitted pipeline and a
machine-readable metrics file, and document the important modeling decisions.
Work independently and run the finished pipeline before you stop.The experiment was not designed as a benchmark for Codex. One run for each condition would not be enough to make a general performance claim. Instead, I wanted to observe which decisions came from the dataset and which came from the project context.
Reproducing the Experiment
To make the experiment accessible, I published the complete project in a public GitHub repository.
The repository contains the 9,000-row synthetic dataset, the data-generation script, the exact prompt used in both runs, both training pipelines, the AGENTS.md file, the prediction contract, the saved metrics, and the figures used in this article. The data is synthetic and contains no real customer information.
The README explains how to install the required packages and rerun each condition. This lets readers inspect each agent's decisions instead of relying only on the metric values shown here.
The final evaluation periods selected by both runs are shown below.
Even though both agents received the same prompt and data, they reserved different months for the final evaluation. This difference also means that the average-precision scores could not be directly compared.
Experiment Result Without Project Context
The result without project context was better than I expected.
The agent inspected the suspicious columns and their relationship with the target. It found that cancellation_date and account_status_30d perfectly revealed churn. The agent correctly identified them as outcome leakage and excluded them from the features.
It also excluded retention_offer_accepted and customer_id.
The agent noticed that the data covered 24 monthly periods. Therefore, it avoided random validation and created expanding-window validation folds. It selected average precision as the main metric, compared several models, and saved a reproducible pipeline.
These were reasonable data-science decisions.
However, the agent still had to decide on information the project did not provide.
It selected September to December 2025 as the untouched test period. The agent reported outreach performance at 5%, 10%, and 20% capacity. It selected the model using the highest mean validation average precision and created another decision threshold based on the F1 score.
None of the decisions were necessarily wrong. However, none came from the project owner either.
The final evaluation included 1,500 rows and 43 churners. The model achieved 0.0836 average precision and captured 12 churners within the monthly top 10% of the risk scores.
The agent produced good code, but it also created part of the prediction problem.
Experiment Result With Project Context
The second agent received the prediction contract through the project instructions.
It reserved July to December 2025 as the final evaluation period, resulting in 2,250 rows and 67 churners. The agent used only pre-July data for model development and maintained time order during validation.
It compared a class-weighted logistic regression baseline with a gradient-boosting model. The prediction contract also provided the selection rule: the challenger could replace the baseline only when it improved average precision without reducing recall at the 10% outreach capacity.
Gradient boosting did not pass the requirement. Therefore, the agent kept the logistic-regression baseline.
The final model achieved an average precision of 0.0748. Within the 10% outreach capacity, the model captured 18 of 67 churners, which resulted in 0.2687 recall.
The agent also saved the prediction moment, target window, excluded columns, validation period, baseline comparison, and operational metric inside the JSON metrics file. The limitation that the model did not estimate the causal effect of a retention offer was also recorded.
The image shows selected information from the actual metrics.json files. The context run did not only use the required decisions during training. It also kept them as machine-readable evidence for future work.
The overall decision comparison is shown below.
The project context did not suddenly make the AI agent better at Python or machine learning. Both runs were already technically capable.
The difference was that the second agent solved the prediction problem defined by the project rather than selecting its own version of the problem.
What Did the Project Context Change
The prediction contract did not contain any model code. It didn't explain logistic regression, tune the hyperparameters, or provide the feature-transformation pipeline.
What it changed was the number of decisions the agent needed to make.
This difference matters because many data-science problems aren't caused by Python errors. The training script might run successfully, the metric might be calculated correctly, and the model file might be saved.
However, the result could still be misleading when the target was defined at the wrong time, future information entered the features, or the evaluation did not represent how the model would be used.
For the churn project, one of the most important sentences was the following:
At the end of every calendar month, the retention team scores customers with an active subscription.
The sentence defined the prediction moment. From this information, the agent could determine what data was available, how to form the target, and why time-based evaluation was required.
Another important sentence was:
The retention team can contact only 10% of customers.
This information turned the task from a generic classification problem into a ranking problem connected with an operational decision.
The value of context engineering was not providing more information in general. It provided the specific information the AI agent could not reliably infer.
What Should Remain as Human Work
Project context could preserve human decisions, but it could not make all the decisions for us.
The data scientist still needed to decide whether churn was the right target, whether the 30-day window matched the business action, and whether the historical customers represented future conditions.
We also needed to decide whether a risk model was the correct method. A model that ranks customers by churn risk does not automatically tell us which customers would respond to a retention offer.
After we defined the decisions, the AI agent could help translate them into the data pipeline, validation process, metrics, and model artifacts.
Conclusion
AI coding agents already understand a lot about Python, Scikit-Learn, and common machine learning workflows. Repeating the same coding explanation in every prompt might not be the most valuable use of project instructions.
More valuable context was the information the code could not explain: when the prediction happened, what information was available, how to simulate the future, which metric represented the operational decision, and what the model result could not prove.
In my experiment, the agent without project context still produced a good pipeline. However, it still had to make several important decisions on its own.
The agent with project context solved the version of the prediction problem that the project had already defined.
For me, this is the role of context engineering in data science.
It does not teach the AI agent how to code the model. It teaches the agent which modeling problem we are trying to solve.
I hope it has helped!
Need More Personal Guidance?
If you are working on your Data Science or AI career and want more personalized guidance, I also offer mentorship through MentorCruise.
We can discuss your career direction, skills, portfolio, technical challenges, or what you should focus on next.







