An EdTech product can perform well in a five-classroom trial and still become a headache once a district buys hundreds or thousands of licences.
The pilot teachers may have received extra training. Vendor staff may have answered questions within hours. Rosters may have been cleaned manually. Everyone knew the trial was being watched. Once the product moves into ordinary classrooms, those favorable conditions disappear.
Teachers are juggling other priorities. New students arrive mid-semester. Passwords stop working. Chromebooks vary in age. A district technology specialist who helped six teachers during the trial may now be supporting six schools.
That gap between a controlled trial and everyday school use explains much of why EdTech pilots fail as procurement tools. A pilot can prove that software works without proving that an organization can use it consistently, affordably, safely, and at the scale being proposed.
There is real money behind this problem. UNESCO’s 2023 Global Education Monitoring Report cited U.S. data indicating that an average of 67% of education software licences were unused, while 98% were not used intensively. The report also cautioned that evidence supporting education technology remains uneven.
Buying more technology is not the same as improving learning. A useful pilot has to test the conditions around the product as seriously as the product itself.
Why EdTech Pilots Fail Before Anyone Logs In
Some weak pilots start with a vendor presentation.
A platform looks impressive. A neighboring district uses it. Funding is available. Senior leaders want to explore AI, adaptive learning, tutoring, assessment, or another priority. The next question becomes: where can we try this?
That sequence makes evaluation harder.
A better process starts with the problem. Digital Promise’s EdTech Procurement Framework places needs analysis and an inventory of existing resources before product discovery. That may sound procedural, but it prevents a costly mistake: buying software first and defining its purpose afterward.
Suppose a district is concerned about Grade 7 mathematics performance.
“Students liked the platform” is not enough to justify expansion.
Before the trial begins, the district should know which students it is trying to help, what teachers currently do, when the product will fit into instruction, how frequently it needs to be used, and what evidence would justify a purchase.
The distinction matters because a product can generate plenty of activity without solving the problem that led the district to shop for software in the first place.
The Trial Is Testing Your School System Too
Schools often treat the technology as the variable being tested. In practice, the surrounding system is being tested at the same time.
A product may have credible research behind it and still be a poor local fit.
Perhaps teachers cannot give it the recommended amount of instructional time. The district’s identity system does not integrate cleanly. Students move between devices that perform differently. The reporting interface requires more administrative work than expected. Accessibility features cover some needs but not others.
A strong product can stumble under those conditions.
The reverse is also possible. An average product can look unusually successful when pilot teachers receive intensive coaching, direct contact with vendor staff, additional planning time, and help that will not exist after district-wide adoption.
Education research offers a useful way to think about the distinction. The U.S. Institute of Education Sciences separates efficacy studies conducted under relatively favorable conditions from effectiveness studies conducted closer to normal practice. Scale-up research asks another question: does the intervention continue to perform across larger populations and more varied settings?
A school pilot is not automatically a formal research study, but the principle transfers well.
If the conditions that made the pilot successful cannot survive expansion, leaders should be cautious about treating the pilot result as proof of scalability.
Real Rollouts Show Where the Friction Appears
Some useful public examples of district EdTech trials are several years old. They should not be used to judge the current versions of the products involved. Their value lies elsewhere: they document what happens when software moves from a sales conversation into working classrooms.
A 2016 project involving Digital Promise and Carnegie Mellon University examined five EdTech products across three Pennsylvania-area school districts. More than 700 students and 30 teachers were involved.
The work found that trials were more useful when districts first identified a need and selected products closely aligned with that need. More focused trials also produced cleaner evidence than experiments scattered across many grades, subjects, and classroom conditions.
The studies recorded less glamorous problems as well: device limitations, technical difficulties, timing issues, and differences in how teachers actually used the tools.
Those are not side issues. They are part of the purchasing decision.
A trial spread across too many variables can produce a pile of data without a clear answer. If results are poor, leaders may not know whether the problem was the software, the devices, the training, the timetable, or the way the trial was designed.
The Mathspace Pilots Show Why Local Conditions Matter

Two Mathspace studies documented by Digital Promise make the point particularly well.
In Rowan-Salisbury Schools in North Carolina, the trial involved 4,197 students and 76 teachers across Grades 6–8. Usage varied substantially between schools and classrooms, and most participating teachers did not reach the product developer’s recommended implementation threshold. Many teachers also reported wanting more help integrating the program into their teaching.
A much smaller Mineola pilot involved 191 fifth-grade students and four teachers. The district had already adopted Mathspace in Grades 6–8, which meant the participating teachers had colleagues familiar with the platform. Teachers reported classroom use of roughly one to two hours per week and said experienced colleagues became an important source of support.
It would be misleading to rank one district against the other. The settings were different.
The useful lesson is that the same product can land in two school systems and encounter very different levels of staff familiarity, support, scheduling flexibility, and classroom adoption.
A product review cannot tell a district whether those local conditions exist. A well-designed pilot can.
Six Problems That Distort EdTech Pilot Results
1. The problem is too vague to measure
“Improve engagement” sounds reasonable until someone has to decide whether the product succeeded.
So do goals such as “modernize teaching” or “use more AI.”
A useful problem statement identifies the people affected, the problem they face, and the point in the workflow where the difficulty appears.
For example:
Grade 7 mathematics teachers need earlier evidence of which students are struggling with fractions because current assessment results arrive too late to adjust instruction during the unit.
Now the product has something specific to address.
2. The pilot team is unusually enthusiastic
Schools naturally recruit willing teachers. That is sensible. The trouble starts when every participant is a technology champion who enjoys troubleshooting new software.
That group may not represent the conditions after a broad rollout.
Include teachers with different levels of technical confidence and pay attention to routine tasks. How long does it take to create classes? Do student rosters sync correctly? What happens when a student changes sections? Can teachers find the information they need without opening three dashboards?
Small frustrations become expensive at scale.
If five pilot teachers each need occasional help from an instructional technology specialist, the burden may be manageable. If 300 teachers need the same help, the staffing calculation changes completely.
3. Training stops at “how to use the platform”
A product walkthrough can show teachers where the buttons are. It does not explain where the tool belongs in instruction.
Teachers need clarity about what the software should replace, what should remain teacher-led, how often students are expected to use it, and how the resulting data should influence teaching.
The Education Endowment Foundation’s implementation guidance organizes change around four broad phases: Explore, Prepare, Deliver, and Sustain. The important word for EdTech teams is sustain.
A launch webinar is not enough if successful use requires teachers to change classroom routines.
Follow-up support becomes more valuable after teachers have used the product long enough to encounter real problems. That is when questions shift from “Where is this feature?” to “How am I supposed to fit this into Tuesday’s lesson?”
4. IT and governance reviews happen after the classroom trial
This is one of the more avoidable procurement mistakes.
Teachers can love a platform that later turns out to create problems with student data, accessibility, account provisioning, device compatibility, reporting, filtering, or security requirements.
Those checks belong before a large classroom trial.
A district should know, at minimum, how users will authenticate, what information the product collects, how data can be accessed or deleted, which browsers and devices are supported, what integrations are required, and whether students with relevant accessibility needs can complete the core activities.
Requirements vary by country and jurisdiction, so schools need to apply their own legal and policy standards. The categories themselves are broadly applicable.
The UK’s Department for Education, for example, treats broadband, cyber security, accessibility, filtering and monitoring, digital governance, devices, networking, cloud services, and IT support as connected parts of school technology planning.
A pilot should not reach week eight before someone discovers a basic infrastructure or governance conflict.
5. Usage data are either worshipped or ignored
A login is not a learning outcome.
But usage still matters.
If the product is supposed to be used three times each week and most students open it twice during the entire month, the final academic results cannot be interpreted without that information.
Define reasonable implementation before launch.
Depending on the product, useful measures might include active students, completed activities, time spent on meaningful tasks, teacher assignment frequency, or sustained use over several weeks.
Then investigate poor adoption rather than simply placing it in a dashboard.
A drop in usage can mean several things. Teachers may lack time. Login problems may be frustrating students. Training may have been too thin. Content may not match the curriculum. The vendor’s recommended dosage may be unrealistic.
Those are different problems and should lead to different decisions.
6. The purchase has effectively been approved in advance
Some pilots are evaluations in name only.
If senior leaders, procurement staff, or the vendor already assume the district will buy the product, disappointing evidence becomes something to explain away rather than something to learn from.
A serious trial needs permission to produce an inconvenient answer.
That answer could be:
- buy and expand;
- use the product only with particular grades or student groups;
- change the implementation model and test again;
- continue the current approach instead;
- stop.
Rejecting a product after a useful trial is not wasted work. Spending heavily on a product that the trial already showed was unsuitable is.
Put the Buying Decision Into the Pilot Plan
Before launch, document what the trial is supposed to decide.
A simple framework is enough:
| Decision Point | Define Before Launch | What It Helps Prevent |
|---|---|---|
| Problem | Specific instructional or operational need | Buying a product without a clear purpose |
| Population | Schools, grades, subjects, and learner groups | Testing with the wrong users |
| Baseline | Current results or workflow | Judging change without context |
| Intended use | Where the product fits into normal practice | Random or inconsistent use |
| Expected use | Realistic frequency or dosage | Misreading low adoption |
| Support | Training, technical help, and ownership | Hidden staffing demands |
| Evidence | Outcomes, usage, and user feedback | Decisions based on one metric |
| Decision rule | Expand, narrow, redesign, retest, or stop | A predetermined purchase |
Do not stop planning at the last day of the trial.
Ask what full adoption would require.
Who creates accounts for new students? Who handles support in September when new teachers arrive? Is professional development included after year one? Can administrators export the data they need? What happens to student information when the contract ends? How much internal staff time will the platform consume?
The annual licence price may be straightforward. The operating cost often is not.
Outcomes-Based Contracting Is Testing a Different Approach
Some districts and organizations are trying to make those responsibilities more explicit in the contract itself.
In March 2026, Digital Promise published findings from an early cohort experimenting with outcomes-based contracting for EdTech. These arrangements can connect parts of vendor payment to agreed outcomes while also specifying implementation responsibilities.
Participating districts reached product dosage requirements for 50% to 95% of students, but the work also exposed practical difficulties: staff time, data access, professional-learning needs, unclear contract terms, and the challenge of translating vendor-recommended usage into ordinary classroom schedules.
Those findings should not be treated as proof that outcomes-based contracts automatically improve student results. The work represents an early implementation model.
Its more useful lesson is that procurement improves when expectations stop being implicit.
Schools do not need an experimental contract structure to apply that lesson. A conventional agreement can still specify training responsibilities, implementation milestones, data access, support expectations, and what happens if usage falls well below the level assumed during procurement.
A More Defensible Rollout Sequence
For most districts, the strongest approach is less ambitious at the beginning.
Start by examining the problem and the tools already being paid for. With evidence of widespread software underuse, adding another licence should not be the default response to every instructional gap.
Next, screen possible products before teachers spend weeks testing them. Review instructional alignment, evidence quality, accessibility, security and privacy, integration requirements, device support, vendor support, and likely long-term cost.
Then choose the smallest trial that can answer the purchasing question.
A district evaluating a middle-school mathematics intervention may learn more from a carefully selected group of classrooms than from launching the product across every middle school immediately.
While the trial runs, watch what is happening rather than waiting for the final survey.
If usage collapses in week three, find out why in week three.
When the evidence is ready, separate three questions:
- Was the product used as intended?
- Did the intended result improve?
- Can the school system reproduce those conditions at a larger scale?
A product can pass the first question and fail the second. It can pass the first two and still fail the third.
That third question is where many expensive rollouts become difficult.
Moving from 10 teachers to 200 is not simply the same pilot with more accounts. It changes support volume, training logistics, communication, rostering, reporting, leadership oversight, and the consequences of downtime.
Scaling deserves its own implementation plan.
What EdTech Vendors Should Learn From Failed Pilots
Vendors also have something to lose when a trial creates unrealistic expectations.
Making the pilot effortless may help close a sale. It becomes counterproductive if the district later discovers that the support provided during the trial was exceptional.
Schools need realistic information about device requirements, integrations, training, staff responsibilities, expected usage, data availability, accessibility, and the level of vendor assistance included after purchase.
The same honesty should apply to recommended dosage.
Sometimes low adoption reflects weak leadership or insufficient training inside the school system. Sometimes the product simply asks too much of a crowded timetable.
If district after district struggles to use a product at the level supposedly required for results, “implementation failure” eventually becomes a product question too.
Software designed for schools has to fit schools as they actually operate.
Final Thoughts
The most useful answer to why EdTech pilots fail is not that schools choose bad technology. The bigger problem is often that the trial tests too little of what will determine success after the purchase.
A procurement team should be suspicious of a pilot that runs only in unusually supportive classrooms, ignores staff workload, postpones technical checks, or cannot describe what would cause the district to say no.
Before scaling any product, ask one practical question: Could we reproduce these conditions across the schools and classrooms where the technology would actually be used?
If the answer is uncertain, buying more licences will not make the evidence stronger.
A good pilot earns the district the right to expand confidently—or to walk away before a difficult rollout becomes an expensive one.





