Self-Learning AI Models
See how learning AI could help contractors automate bidding, vendor management, and other complex work.
Matt Wolfe
Co-founder, Bidlo
AI is beginning to learn from experience, which could make it far more useful for everyday construction work. This article explains how these advances could help contractors automate bidding, vendor management, and other complex tasks—and how Bidlo is preparing to put them to work.
Today, we'll discuss AI models that actually learn, where the industry is now, where it's headed, and what you can do today to prepare. The best place to start is how these models are built. Take GPT 5.6, OpenAI's most recent model.
The GPT-5 in the name represents the base model. To build such a base model, you feed a massive amount of data and compute into a system that outputs weights. These weights mathematically represent the model's understanding of the world.
In simple terms, these weights help the model match new inputs to what it has already learned so it can provide answers. The 0.6 in GPT 5.6 comes from fine-tuning and reinforcement learning (RL). Labs like OpenAI or Anthropic create pass-fail tests and repeatedly have the model perform and readjust to optimize for accuracy.
So, building models involves spending a lot of money and resources to generate the base model, then applying these pass-fail tests to fine-tune and improve performance on specific tasks. That's how you get GPT 5.6. The challenge is the enormous cost and resources required, which explains the long gaps between versions, such as between GPT 5.5 and 5.6, and even longer between GPT 4 and GPT 5.
These models are hitting multiple walls.
First is the data wall. Much of the data used to train these models is public, but a lot of remaining data is private, such as company data, or it exists in poor or non-digital formats—for example, public construction or Department of Transportation data is often inaccessible or unstructured. This makes training with such data very difficult.
A bigger wall is the training method itself. These models excel on clear pass-fail online tasks—like writing code and checking for errors—but real-world applications such as running a business or navigating complex situations are messy and ambiguous. This messiness makes it hard to train models using strict pass-fail approaches, which explains why autonomous driving has taken so long to develop compared to other AI models.
Finally, even when it’s possible to retrain or use RL on messy real-world data, retraining is computationally expensive and time-consuming, making constant learning impractical. Humans excel at this process—we quickly absorb new information, reconsider our approach, and improve.
This process, called sample efficiency, is where models currently lag because retraining and fine-tuning take so long, limiting their ability to handle real-world tasks effectively. So far, the strategy has been to train models on as much data as possible, hoping they have a rough idea for any scenario. Humans have less breadth of knowledge but can adapt quickly in new situations.
The future of this technology lies in models that can learn from ongoing interactions and adjust their weights accordingly. Many labs focus on implementing this ability with greater sample efficiency. Temporary solutions you may have seen include memory systems or agent harnesses that save rules or memories, but these are band-aids—they do not represent true learning but rather give the model a fixed set of rules to follow.
Promising technologies to watch include On-Policy Self-Distillation (OPSD). This technique allows a model to adjust only the weights related to a specific task after receiving feedback, rather than retraining the entire large model, reducing time and energy costs.
Another gaining traction is the concept of “dreaming.” Neuroscience suggests humans dream to simulate scenarios and prepare for future situations. Similarly, models can simulate many variations of an experience to find optimized responses and adjust their weights accordingly. This limits downtime and enables autonomous reinforcement learning.
Meanwhile, open-weight models are catching up to larger lab models. For example, with sufficient hardware, an organization could run an open-weight model and use reinforcement learning techniques to teach it how to operate within their specific business context. This ability to own, train, and improve a model on proprietary data is the ultimate goal.
What are we advising customers to do now, and how is Bidlo preparing? A model needs a "classroom" where it can learn your business operations, understand right and wrong, work alongside you on complex tasks that might take days or weeks, then eventually perform autonomously. These will be your digital employees of the future.
As you see updates to the Bidlo platform in coming months, keep in mind this is our direction. Today, the platform helps with procurement, bidding, solicitation assembly, vendor management, schedule management, and various AI-driven automation tools. The end goal is not for you to manually step through each task but to delegate them to an agent that continually learns and optimizes to improve your bottom line.
