【
Smart CityOnline Brand Column】If someone is swimming in the reservoir, please alert them in a timely manner. "In the past, this was an instruction given to safety administrators, but at the 2026 World Artificial Intelligence Conference (WAIC 2026), Hikvision showcased a new possibility - assigning this task directly to a camera.

On July 17th, the first day of WAIC 2026's opening, Hikvision released a multimodal intelligent camera, achieving the goal of "making the camera work with just one sentence". For example, in the scenario of reservoir safety management, traditional cameras see "people" and "water bodies", while multi-modal intelligent cameras with pre-set tasks can not only monitor risk areas in real time, accurately describe people, states, and environments in the scene, but also understand the scene, judge "drowning risks of personnel swimming", and actively link horns to issue warnings. This ability to understand, comprehend, and execute signifies that cameras are evolving into proactive "intelligent assistants".

This breakthrough benefits from the deep integration of multimodal large models and edge computing architecture. The relevant person in charge of Hikvision introduced: "We have installed a multimodal large model into this compact body, allowing it to not only perceive accurately, but also quickly recognize, judge, and respond, adapting to scenarios such as near water reminders that are related to safety and require high real-time performance. Compared with traditional cameras, multimodal intelligent cameras have stronger scene understanding ability, more flexible task configuration methods, and more real-time edge processing capabilities.
Specifically, traditional cameras can only see what is in the picture, while cameras with a single visual model can understand what is in the picture, but cannot respond to human language commands. And multimodal intelligent cameras have achieved a triple leap through multimodal capabilities: firstly, interactive upgrading, from writing complex rules to directly stating requirements; The second is to upgrade understanding, from identifying "a group of people" to understanding "they are swimming and there is a risk"; The third is the upgrade of decision-making, from passive alarm to active linkage of sound and light, notification and control.
In terms of computing power and parameters, the multimodal intelligent camera integrates 26T computing power, built-in 2B parameter quantity multimodal large model, supports deep understanding of graphics and text and semantic analysis, and can achieve multimodal interactive closed-loop on the end side. It can be widely applied in fields such as public safety, urban governance, industrial manufacturing, and safety production.
It is worth mentioning that multimodal intelligent cameras support natural language deployment, which can effectively improve business efficiency. Taking the application in safety production scenarios as an example, users only need to input a natural language, such as "If personnel enter the construction area without wearing reflective clothing, please call the police". The system immediately converts it into dynamic control rules. Once the real-time video stream on the end matches the target, it will immediately trigger an alarm and support the linkage of surrounding devices such as sound and light, access control, etc. From passive retrieval to active defense, from post verification to in-process intervention, achieving closed-loop management.
At the WAIC 2026 Hikvision exhibition area, visitors can also personally experience this "within reach" intelligence. By clicking on the entry on the screen, such as "What scene is this, analyze the risk behavior in the scene" or "Detect whether someone is swimming in the picture, if so, give an alarm", the multimodal intelligent camera can automatically analyze the video content and generate real-time graphic and text Q&A results, warning information on the screen, and link to issue a voice alarm prompt "Swimming is prohibited here, please leave as soon as possible".
This demonstration not only visually demonstrated Hikvision's strength in image and text understanding and natural language interaction, but also allowed the live audience to personally feel that the large model is not only "understandable", but also "conversational" and "functional", and is a truly intelligent assistant that can be rooted in frontline scenes.
From reservoir proximity reminders to factory safety production, from urban governance to smart retail... In the past, many intelligent analysis scenarios were limited by network bandwidth or cloud computing costs, making it difficult to achieve real-time response. Hikvision, with its profound accumulation in the field of intelligent IoT, continues to promote the landing of AI technology at the edge. The release of multimodal intelligent cameras now marks the true sinking of AI capabilities to the last mile of industry applications, injecting surging new momentum into the digital transformation of thousands of industries.