Using a smartphone at a construction site isn’t as easy as you might think. Both hands are full of tools and materials, and fingers wearing safety gloves often fail to register touch inputs properly. In basements, tunnels, and inside concrete structures, even the cellular signal becomes unreliable.
No matter how advanced internet-connected AI becomes, it’s difficult for it to provide practical help if it can’t be used on-site—where record-keeping is most critical.
Digital Presso Co., Ltd. has been selected for a national R&D project aimed at addressing these on-site limitations and is currently conducting research.
Through the 2026 Startup Growth Technology Development Project (Ddimdol) Strategic Technology R&D program, supported by the Ministry of SMEs and Startups, the company is developing voice-based on-site documentation technology that operates on a smartphone even in environments without an internet connection.
In this article, we’ll explore why this technology is needed, how it differs from existing generative AI, and how it can be utilized in real-world settings.
The Real Reason Field Workers Can’t Use Apps
Existing field work systems—whether mobile apps or web-based—are designed on the premise that workers will directly interact with the screen, but in real-world field settings, this premise often does not hold true.
During work, both hands are occupied with tools and materials, and protective gloves make it difficult to use a smartphone’s touchscreen. Inside underground spaces, tunnels, or structures, cellular reception can be unstable or cut off entirely.
As a result, on-site documentation gets put off. The moment a worker thinks, “I’ll sort this out when I get back to the office,” the context of the photos becomes blurred, and inspection details end up relying on memory. Reports naturally get delayed as well.
These missing or inaccurate records become a problem later on when quality and safety must be verified. This is because, precisely when evidence is needed, there are not enough usable records left.
Where Digital Transformation Stalls on-Site
For a long time, digital transformation in the construction industry was viewed as simply “a matter of creating better apps.”
However, the difficulty in using apps on-site often stems not from a lack of features, but from the fact that workers’ hands are occupied and they lack a stable internet connection. If the problem lies not in the software’s sophistication but in the input methods and connectivity environment, the solution must start there.
“Take a photo of the shotcrete surface in Tunnel Section 3 and log the inspection.”
The technology being developed for this project understands sentences spoken by workers in their everyday language and automatically executes the corresponding app functions.
For example, when a worker says, “Take a photo of the shotcrete surface in Tunnel Section 3 and log the inspection,” the camera launches, and the captured photo and inspection details are automatically linked to that specific section and work process.
Previously, workers had to manually open the screen, navigate to the menu, select an item, enter the details, and save them; from now on, however, all these steps can be completed with a single voice command.
This applies to over 40 types of repetitive on-site tasks, such as photography, record-keeping, inspections, reporting, regulatory lookups, and manual searches, and all processing takes place within the smartphone without an internet connection. This also means that the worker’s voice commands and on-site data are not transmitted to external servers.

It works even without an internet connection
Generative AI services like ChatGPT and Claude are structured to send user input to external servers for processing and then return the results. In other words, even if they offer excellent performance, they cannot be used if the internet connection is lost—and field documentation and inspections are particularly critical in locations where communication frequently drops out, such as underground, in tunnels, or inside structures.
Therefore, this technology is designed to embed the AI model directly within the smartphone device, enabling it to interpret voice commands and execute necessary functions regardless of network connectivity.
It performs actual tasks, not just providing answers
Existing generative AI and voice recording apps focus on converting the user’s speech into text or generating responses that match the request. Afterward, the user must still manually determine which item at which site to associate the generated content with.
In contrast, this technology goes beyond simply understanding speech; it directly executes app functions that match the user’s intent. For example, if you say, “Take a photo of the plumbing on the third floor and log the inspection,” the camera launches, and the photo and inspection details are saved to the corresponding task item.
In other words, the system is designed so that the user’s words lead directly to actual on-site tasks, rather than being treated as mere text.
Understands Construction Site Jargon
Construction sites involve a wide range of expressions, including everyday language, specialized construction terminology, abbreviations, and site-specific terms for each type of work. When combined with workers’ accents and regional dialects, it can be difficult for general-purpose voice AI to accurately grasp the meaning of what is being said.
In this project, we are training the AI primarily on Korean speech used in actual construction sites. We plan to first improve recognition performance for construction-specific terminology and on-site expressions, and then gradually expand the scope of training to include regional accents and dialects.
Site data does not leave the premises
Since the entire process of interpreting and processing voice commands takes place within the device, sensitive on-site information is not transmitted to external servers.
In public infrastructure and security sites where handling security-critical data—such as blueprints, equipment information, and on-site conditions—is essential, this processing method can be a key factor in deciding whether to adopt the technology. Since there is no need to build or operate a separate AI server, this offers benefits not only in terms of data security but also in terms of operational costs.
No separate dedicated equipment is required
Dedicated wearable devices that enable voice-based task management are already commercially available. However, some products rely on external servers for Korean voice processing, and with prices exceeding $3,000 per device, the cost burden is too high for widespread adoption across an entire site.
This technology is being developed so that workers can use it directly on the smartphones they already use. A key feature is that there is no need to purchase separate equipment, and voice-based functions can be applied while maintaining a familiar work environment.
Key Differences
Generative AI Services: Operate via an internet connection and primarily provide text-based responses
Overseas Field Management Services: Use of key AI features is limited when communication is unstable or interrupted
Voice Recording and Dictation Apps: Convert speech to text, but users must manually categorize and save the recordings
Dedicated Wearable Devices: Require the purchase of separate equipment, creating a financial burden in terms of Korean speech processing and implementation costs
Digital Presso’s Technology: Processes Korean speech directly within a smartphone and executes field tasks according to the user’s intent
The core of this technology lies in implementing Korean speech recognition, on-device processing, flexible voice commands, and on-site task execution within a single framework. Within the scope of our own research, we found it difficult to identify any field-specific technology that meets all four of these criteria.

Construction Management and Safety Inspections
This technology can be applied first and foremost to daily, repetitive record-keeping tasks on-site. Key application areas include construction management, maintenance, safety inspections, public facility operations, and power/equipment management.
Even in situations where it is difficult to use both hands—such as while on a ladder or in front of a distribution panel—or when wearing protective gear, users can record photos and inspection results using only their voice. This allows for more accurate and systematic documentation of details that are easily overlooked on-site, without disrupting the workflow.
Facility Inspection and Maintenance After Completion
It can be utilized not only during the construction phase but also for facility maintenance after completion.
Buildings and equipment are used for long periods after completion, and regular inspections and repairs are continuously carried out during that time. Even in environments where it is difficult to use your hands freely—such as in front of pipe valves or on rooftop equipment—you can easily record inspection results and corrective actions using only voice commands.
Expansion into the Power, Public Utilities, and Facilities Management Sectors
The technology’s architecture is not limited to specific trades or industries. By configuring the tasks to be performed via voice and the command system by sector, it can be applied to a wide range of industrial environments.
Starting with construction terminology and tasks at building sites, the system can be expanded to adjacent fields such as power, public infrastructure, and facility management. In the future, by adapting the language and workflow systems to local environments, it can also be utilized at overseas infrastructure maintenance sites.
Any Work Environment Where Hands Are Not Free
The core condition this technology aims to address extends beyond the construction industry itself to encompass all environments where workers cannot use their hands freely while working.
The same technical framework can be applied in environments where both hands are occupied with tasks—such as welding, precision work, or wearing protective gear—and where communication is restricted. We are developing the technology with scalability in mind so that it can eventually be utilized in smart glasses and helmet-mounted wearable devices.
Small AI Fitting Inside a Smartphone
To run AI on a smartphone, the model must be optimized for limited memory and computational power. For this project, we are utilizing SLM (Small Language Model), a compact language model designed to run on smartphones.
However, simply reducing the model’s size does not make it ready for immediate use in the field. The accuracy of command interpretation may decrease during the optimization process, and if the model is not sufficiently optimized for the device, responses may be delayed, or the device may execute functions that do not match the user’s intent.
Therefore, the core of this project is not simply to implement a small AI. The goal is to quickly and accurately understand the user’s voice commands even within the limited environment of a smartphone and link them to the appropriate on-site tasks.
Development Goals Set for This Project
Based on our own measurements of publicly available lightweight models applied to actual devices, we aim for the following relative improvements:
Reduce the time between the end of a voice command and the execution of the function by approximately 70%
Approximately 30% improvement in the accuracy with which functions are executed according to the user’s stated intent
Output structured data that the app can process immediately
Minimize memory usage to a level that does not interfere with everyday smartphone use
This research goes beyond simply developing new technology as a standalone product. The goal is to integrate voice-based input into Digital Presso’s field management platform, RenameDP, and implement it as a feature that can be utilized in real-world work scenarios.
RenameDP is a platform that systematically accumulates records generated at general construction and electrical construction sites and automatically organizes them into necessary documents. Location and time information are automatically linked to captured photos, and AI analyzes the photos to generate construction records. Communication at each site is compiled into daily work reports, and risk assessments and TBM records can also be managed as historical data and supporting documentation.
The more than 40 types of tasks addressed in this project were also selected based on RenameDP’s actual features and on-site workflows, rather than on hypothetical scenarios.
While the existing RenameDP is a platform for creating and managing site records, this research adds the capability to perform record-keeping and tasks hands-free. We aim to lower the input barrier between the field and digital systems by enabling workers to execute necessary functions using only their voice, without having to remove gloves or navigate multiple screens. In this process, wearable devices such as smart glasses are used to collect photo and voice data.
We plan to integrate the developed technology into RenameDP rather than releasing it as a separate product. Even after the R&D phase, we intend to continue utilizing it in actual field settings to verify its performance and usability.

A Digital Presso representative stated, “The reason records fall behind on-site is not because workers are negligent, but because the current system requires them to remove their gloves and interact with the screen multiple times to create a record,” adding, “The goal of this project is to enable users to complete photo capture and inspection logs using only their voice within a smartphone, even in areas without internet connectivity.”
Apps that require workers to remove their gloves to operate, or AI that relies on an internet connection, have clear limitations when used in the field.
AI that operates within the device itself is one way to address these limitations. It can be used even when communication is lost, and since it does not transmit voice or field data to external servers, it can be applied in environments where security is critical.
However, achieving both fast processing speeds and high accuracy within the device’s limited performance capabilities is no easy task. The figures presented in the text are not current achievements but rather R&D goals that must be validated by the end of the project.
A Digital Presso representative stated, “We plan to implement the technology validated in this project as a feature of RenameDP rather than as a separate product,” adding, “Ultimately, we aim to create a work environment where workers do not have to stop what they are doing to keep records.”
This content was produced by Digital Presso Co., Ltd.
The improvement figures presented in the text are relative improvement targets set based on results measured in-house by applying publicly available lightweight AI models to actual devices, and they represent the goals to be achieved by the end of the project.