The research question
How can a model connect observations to the right asset, follow changes over time and recognise when evidence is stale, incomplete or contradictory? We study multimodal context, task memory and uncertainty.
Welf Labs / Research agenda
Industrial AI needs to understand the situation, adapt when conditions change and learn from experience. Welf Labs focuses on situational awareness, adaptive intelligence, self-improving AI models and specialised language models for industrial work.

Situational awareness
A machine alarm, an inspection image and yesterday’s work order tell different parts of the story. Our research asks how AI can combine them into an accurate view of what is happening now.
How can a model connect observations to the right asset, follow changes over time and recognise when evidence is stale, incomplete or contradictory? We study multimodal context, task memory and uncertainty.
Maintenance triage in energy and process plants, aerospace service preparation, maritime machinery inspection and manufacturing changeover checks. The useful output is a supported account of the situation and the information still needed.
Reconstruct operating states from timestamped records and images. Measure missed changes, unsupported conclusions, source accuracy and the time an operator needs to review the result.
Adaptive intelligence
A new part variant, an unavailable tool or a revised instruction can break a workflow that worked yesterday. Adaptive intelligence concerns how the system responds during the task.
Explore robotics and physical AIResearch how agents recognise a change, select useful context and revise the next step. Compare fixed workflows with systems that can replan, use another permitted tool or ask an operator for help.
Investigate robotic handling across part variants, manufacturing changeovers, shipping exceptions and maintenance planning. Study which skills transfer between sites and which need additional examples.
Test unfamiliar layouts, missing inputs, tool failures and changed priorities. Compare completion rate, recovery time, repeated actions, operator interventions and the cost of adapting.
Self-improving AI models
An expert corrects an answer. A completed task reveals a better method. We study how that experience can improve the next model while preserving what it already does well.
Explore SOP automationCompare fine-tuning on reviewed corrections, preference learning, distillation and reinforcement learning where outcomes can be reliably scored. Study how repeated updates affect earlier capabilities.
Technical-document extraction, maintenance recommendations, recurring SOP exceptions and robot skills learned from demonstrations and operator interventions.
Keep unseen cases and earlier tasks outside the learning loop. Measure improvement, regression and review effort across updates. Compare changes to model weights separately from changes to retrieval, prompts or tools.
Specialised industrial LLMs
Part numbers, tolerances, material grades and procedure revisions carry meaning that general answers can miss. Our agenda examines how language models acquire that knowledge and apply it to a defined task.
Explore private-cloud AICompare retrieval, fine-tuning and their combination. Investigate smaller models, German and English technical language, and multimodal extensions for drawings and tables. Assess continued pre-training where the data and results justify it.
Extract RFQ requirements, reconcile specification changes and interpret material certificates. Test units, tolerances, component identifiers and source revisions across automotive, manufacturing and metals workflows.
Retrieve asset-specific maintenance procedures in energy, oil and gas, petrochemicals, aerospace and maritime operations. Investigate SOP retrieval and deviation-evidence preparation in pharmaceuticals.
Compare domain accuracy, source support, reviewer corrections, latency and cost per accepted result. Test on document families the model has not seen, including missing information and conflicting revisions.
A shared research method
Each research question needs a defined task, a credible baseline and evidence that survives a change in conditions.
Use authorised industrial examples with domain review. Separate development data from unseen assets, document families, sites or time periods.
Compare methods with equivalent information, tools and budgets. Report quality alongside review effort, recovery time and operating cost, including the difficult cases.
The intended outputs are technical reports, evaluation sets and reference implementations. Record the conditions under which a result holds so a delivery team can assess it for its own workflow.

Engineering note
How to build a representative test set, measure results against the current process and decide when a workflow is ready for more automation.
Read the noteA difficult operating condition, a specialised dataset or a recurring failure can be the starting point. Tell us what the system needs to understand and how you would judge a useful result.
Start a conversation