Everyone is talking about prompt engineering: which model is best, which IDE has the smartest autocomplete, which $200/month plan gives you the most tokens.
We think most of this misses the point.
The idea behind everything we build at Alpha Bits is a way of thinking we call Data-First Principle Thinking: the belief that raw data, properly collected and structured, will lead you to conclusions that speculation alone won't.
Scientists call it the empirical method, and analysts call it evidence-based decision-making. We gave it a name that fits how we work across software, hardware, energy research and AI, and we built a company around it.
Tool fatigue is a symptom
If you're a developer or business leader in 2026, you're drowning in AI tools. New models come out weekly, IDEs become "AI-powered" overnight, and every vendor promises that its platform will transform your workflow.
Yet most teams using these tools still get mediocre results, because they skip the foundation. They give AI tools unstructured context, vague instructions and scattered data, then blame the model when the output is generic.
Useful output from AI comes mostly from the context you give it, much more than from the wording of the prompt.
We think of this as a shift from prompt engineering to context engineering. The tools will change every quarter. Your internal knowledge, data structures and documented decisions build up and become more valuable over time.
The data pyramid
We use this framework for every project, experiment and business decision at Alpha Bits:

The pyramid is simple, but moving up it takes work. Here's how it plays out on an IoT monitoring project.
At layer 1, raw data, the team needs to monitor energy usage across several sites. Before writing any code, they list the exact fields required: timestamp, site_id, sensor_type, temperature_celsius, power_draw_watts, battery_voltage, device_firmware_version. That's seven specific fields, chosen to answer the question "Is our system losing efficiency overnight?", instead of "some sensor data."
At layer 2, accessible context, the raw sensor readings are flowing into a database, but they're messy. Different sensors use different units, and some devices report every 5 seconds while others report every 30. The team builds a Data Dictionary that standardises everything: temperatures in Celsius, power in watts, timestamps in UTC. They set up a central InfluxDB with Grafana dashboards. Now anyone, human or AI, can query the data without first spending two hours working out what temp_3 means.
At layer 3, hypothesis and the "AI magic," the team has clean, labelled data in one place and can ask the real question: "Which sites are losing more than 5% thermal efficiency during overnight discharge cycles, and does the pattern correlate with ambient temperature or firmware version?" Give an AI tool that hypothesis plus the structured data and you get a specific answer you can act on, instead of a generic essay about thermal efficiency.
Getting through these layers takes weeks. But every later project that touches the same data inherits the structure you built.
Level 1: Raw data (the base)
Before you touch an AI tool, ask: What raw data fields are required to make a conclusion?
That sounds obvious, but most teams skip this step. They start analysing whatever data happens to be convenient instead of working out what data is necessary.
In practice, this means a few things for us. Before building an IoT monitoring system, we list every sensor reading we need: temperature, humidity, power draw, timestamps, device IDs and location tags, as specific fields. Before optimising a client's operations, we audit their existing data sources; if they have five different POS systems across 200 outlets, we map every field in every system before writing any integration code. Before debugging a production incident, we collect logs, error traces, network captures and timeline data before forming hypotheses about what went wrong.
If you don't have the raw data, you can't draw conclusions, and no AI model makes up for missing data. Document what you need first, collect it second and analyse it third.
Level 2: Accessible context (the middle)
Raw data is useless if it's scattered across fifteen spreadsheets, three messaging apps and someone's memory. The second level is making your data clean, properly labelled and accessible from one central location.
At Alpha Bits, we enforce this with a few habits.
Every project gets a Data Dictionary, a living document that defines each data field: its source, format, update frequency and owner. When an AI tool needs context about our codebase, our infrastructure or a client's business logic, this document is the first thing it reads.
We write down every SOP, architecture decision and meeting outcome in a shared, searchable place. Slack threads and email chains don't count. It has to be a document that AI tools can access and reference.
We name things precisely. A database column called temp could mean temperature, temporary or template, and we don't allow that ambiguity. Names like sensor_temperature_celsius, is_temporary and email_template_id help your future self and every AI tool that will ever read your schema.
This level is tedious, and most teams skip it because it feels like overhead. We've learned the hard way that every hour spent organising data saves ten hours of confused AI output and debugging later.
Level 3: Hypothesis and validation (the peak)
Only once the base and middle layers are solid do you reach the part most people start with: asking AI to do something useful.
At this level, you write clear problem statements or hypotheses that can be checked against your actual data:
- "Based on last quarter's sensor readings, is our thermal storage system losing more than 5% efficiency during overnight cycles?"
- "Which of our client's 200 outlets has the highest variance between POS-reported revenue and bank settlement amounts?"
- "Given the last 90 days of Node-RED flow logs, which automation sequences have the highest failure rate?"
These are testable hypotheses grounded in structured data, so the AI's job is to process the data and surface patterns instead of guessing what you meant.
The AI isn't doing anything special here. The context you've provided makes a generic answer hard to give, so it answers with specifics.
Real examples of DFPT in practice
The accidental invention
In late 2022, one of our team started a thermal energy storage experiment: a sand battery prototype on an apartment balcony in Saigon. On paper, this person had no business doing energy research. Their background was software, mostly SQL queries, ERP systems and data pipelines.
They approached it the way they'd approach a database problem, with data first.
Every experiment had IoT sensors, and every temperature reading, power input and heating cycle was logged with timestamps to a local database. They didn't know what they were looking for, but they expected that the raw data would eventually reveal something.
It did. Within months, patterns in the data pointed to a thermal isolation technique that nobody in the existing literature had documented, and that technique became a US patent. It came from methodical data collection that made the pattern visible, not from genius or luck.
You don't need to be a domain expert to find something new. You need to collect the right data and pay attention to what it shows you.
The security incident
A production system was hit by malware, an "I Love You" variant, at the start of a week-long holiday. The team wasn't available, and the founder was handling it alone.
The usual response would be panic, guesswork or waiting until Monday. Instead, the founder applied DFPT and collected the raw data first.
Within hours, with AI tools processing hundreds of thousands of log lines, the attack vector was identified, the compromised entry point was pinpointed with exact timestamps, and a patch was deployed. The AI didn't "know" what had happened. Given properly structured server logs, network captures and access records, it found the needle in the haystack quickly.
The lesson is about logs more than AI: structured, centralised logs make any analysis tool, AI or human, effective when it matters.
The F&B data unification
A multi-brand food and beverage chain with 200+ outlets across a Southeast Asian country came to us with a familiar problem: data everywhere and insight nowhere.
They had five different POS systems and manual daily reports from store managers. Revenue reconciliation happened in spreadsheets that hadn't been audited in months. The CEO's dashboard showed one number, the CFO's showed another, and the store managers' daily reports showed a third.
Most consultants would have started with a tool recommendation, such as "You need Tableau" or "Let's implement Power BI." We started with the Data Dictionary.
For three weeks we did nothing but map every data field across every system: what each POS called a "transaction," how each outlet reported "daily revenue," and which fields were auto-generated, manually entered or calculated. The taxonomy alone filled a 40-page document.
Once the data was clean and in one place, the insights were obvious. The first round needed no AI, because the discrepancies were visible to the naked eye. AI came in for the second round: pattern detection across 200 outlets, anomaly flagging and predictive inventory recommendations.
Before you automate analysis, understand what you're analysing. The Data Dictionary is the most boring and most useful document we produce.
The AI team simulation
Recently we built an interactive thermal simulation, a complex physics visualisation that lets users explore how sand batteries charge and discharge. An entirely AI-powered team finished it in a few hours. NotebookLM acted as lead researcher, consolidating years of past experimental data and research notes into a structured brief. Google Stitch was the UI designer and generated the interface components. Antigravity was the full-stack developer and wrote the simulation logic.
This only worked because years of DFPT habits had already structured the raw data. The research notes were documented, the experimental results were in databases, and the physics parameters were catalogued in a Data Dictionary.
Without that foundation, the AI team would have produced a generic, inaccurate simulation. With it, they produced something that domain experts validated as technically accurate.
AI teams are only as good as the data context they inherit, and the structured data you build now feeds every future project that touches it.
How to start thinking data-first
If you want to try this, here's where to begin.
0. Start with pen and paper
I know I just spent 2,000 words on data structures and AI context, but don't open a single app yet.
Get a notebook, any notebook. For one week, write your daily journal and to-do list by hand: what you worked on, what decisions you made and what's blocking you.
If pen and paper feels too analogue, use the Notes app on your phone. The tool doesn't matter. What matters is that you do it every day without exception.
After a month you'll notice a change, because the habit of writing things down forces you to separate what actually happened from what you think happened. DFPT lives in that gap.
The Data Dictionaries and centralised documentation come later. You can't structure data you never bothered to capture, so the habit comes first.
1. Then start a Data Dictionary
Pick one project, system or workflow. Document every data field: what it's called, where it comes from, how often it updates and who owns it. This one document will change how you, and any AI tool, work with that system.
2. Track everything in writing
Write down every SOP, feature decision and meeting action item in a shared, searchable place, as documents rather than chat messages or notes in your head.
This is how you build the context layer that makes AI useful. The next time you need an AI tool to help with a decision, you'll have six months of documented context to give it instead of a vague prompt from memory.
3. Define hypotheses before opening AI tools
Before you type a prompt, write down what you're trying to learn. What specific question are you answering? What data would prove or disprove your hypothesis? What format do you want the answer in?
Doing this three-sentence exercise before you use any AI will improve your output more than any model upgrade or subscription tier.
4. Audit your data gaps
The most useful thing you can discover is what data you're not collecting. Run a gap analysis: for every decision you make regularly, ask "What data am I using to make this decision? Is it complete? Is it current? Is it accessible?"
The gaps you find will show you where to invest next.
Context engineering vs. prompt engineering
The industry is obsessed with prompt engineering, the craft of writing the perfect instruction to get better output from a model. Better prompts do help, but they hit a ceiling. Context engineering keeps paying off the more you put into it.
A well-crafted prompt with no data context still produces a generic answer. A rough prompt with rich, structured data context produces a specific, useful one. We've seen this happen hundreds of times across projects.
That's why we spend more time organising our data, documenting decisions and building Data Dictionaries than tweaking prompts. The context does most of the work, and the prompt points the AI in the right direction.
Why we work this way
For us, Data-First Principle Thinking is as much about honesty as about data.
It means saying "I don't know yet, let me look at the data" instead of "I think I know, let me find data that confirms it." It means writing things down even when no one is checking, and letting the numbers contradict your assumptions when they do.
We built Alpha Bits on this. It led us to inventions we had no business making, insights we didn't have the expertise to predict, and products that work because they're grounded in what the data says.
The tools and models will keep changing, but the habit of collecting, structuring and reasoning from data won't go out of date. Start there.
Data-First Principle Thinking is how we approach everything at Alpha Bits, from thermal energy research to AI-powered development. We write about it because we think the framework is more useful when it's shared.