Skip to main content

Fusion tracker

The Fusion Tracker module works to fuse the output from several more specialized trackers into a single more comprehensive and accurate tracking data output. The module is also responsible for incorporating other features such as, for example, producing object snapshots and producing object geographical coordinates.

The data processing pipeline that includes Fusion Tracker instance can be visualized as,

A single Fusion Tracker module instance is dependent on data from module instances of the types:

Since this module builds on data from previous modules, the configuration of the input modules will also effect the fusion module instance.

Instances​

This module is generally pre-configured with one instance per physical video channel.

Notes:

Function​

The Fusion Tracker module tracks and estimates state of objects in the scene. The estimated state of each tracked object is continually updated to provide up to date estimates for each object's state currently present in the scene.

  • Update frequency, approx 10Hz – on multidirectional cameras when multiple instances are producing data the framerate may be lower.
  • Latency, approx 1s – the latency of the output varies but is generally around 1 second.
  • Heartbeat, every 2s – If no objects are detected in the scene a heartbeat message (an empty scene message) will be sent every 2s.

To indicate that an object's track has been terminated (never to be updated again) a delete track operation is sent. Tracked objects that are classified humans may be re-identified (ReID), this is represented by sending a rename track operation saying that an id has been renamed. For more information on ReID, refer to the description of ReID on the ReID concept page.

Object state​

Depending on how the module is configured, what sensors are considered, and other factors the exact estimated state might be different. An overview of the object state estimated for each object is found below.

General:

  • Class - Object classification.
    • Human
    • Human Face
    • Vehicle - Car
    • Vehicle - Bus
    • Vehicle - Truck
    • Bike
    • License Plate
  • Vehicle Color - Vehicle color
  • Human Upper Clothing Color - Human upper clothing color
  • Human Lower Clothing Color - Human lower clothing color
  • License Plate Text - License plate text (requires license plate text feature)
  • Image - The current snapshot image of the object (requires Object Snapshot feature)

2D Image Perspective:

  • Bounding Box - Object position as bounding box in the image
  • CenterOfGravity/Centroid - Position of the object as a point in the image

3D Device Perspective:

3D World Perspective:

Track operations​

Dewarped views​

On fisheye cameras with an ARTPEC-8 or ARTPEC-9 SoC, the Fusion Tracker module runs only on the 360° overview. You can still retrieve data for objects visible in the dewarped views: Panorama, View Area 1-4, Corner Left, and Corner Right. In these views, the objects' bounding box coordinates are transformed to align with the respective video view. Note that, due to an object's orientation in the 360° overview, its bounding box may become larger after transformation to a dewarped view.

Be aware of limitations when consuming data from multiple channels.

Output protocols​

This module's output can be retrieved using a variety of methods. Depending on what method is used the data can be sent on different formats, below is a full list of ways to receive the output.

ProtocolName/AddressFormatGuide
RTSPAddress: rtsp://ip-address/axis-media/media.amp?analytics=polygon Source: AnalyticsSceneDescriptionONVIF tt:FrameConfigure scene metadata over RTSP
MQTTcom.axis.scene.frame.v1ADF Framescene metadata over MQTT
MQTTcom.axis.scene.object_snapshot.v1ADF Object Snapshotscene metadata over MQTT
Device Data Hubcom.axis.analytics_scene_description.v0.betaADF Beta FrameACAP example (consume-scene-metadata)

Configuration​

There are a number of configuration options effecting the output from this module. All of the below configuration effects all instances unless otherwise stated.

Some configuration options require a device restart to take effect. Best practice is to restart the device when a configuration options is changed to be sure it has taken effect.

All configuration options is not available on all devices.

DescriptionMethodDefault valueNoteConfiguration guide
Object Snapshot featureRest API, http://<servername>/config/rest/object-snapshot/v1/enabledOFFSee details below.Enable and Retrieve Object Snapshots
Geographic coordinates featureConfigure device geolocation and device geoorentation using designated cgi:sOFFThis feature is implicitly enabled by configuring the device geolocation and geoorientation. Once enabled it can not be turned off other than by a device reset. Only available on some devices, see more details below.Enable Geographic Coordinates
License plate text featureInstall AXIS License Plate Verifier ACAPOFFThis feature is implicitly enabled by installing the AXIS License Plate Verifier ACAP. Only available on some devices, see more details below.Install ACAP
Image rotationImage rotation CGIDevice dependentThe object state related using a 2D image perspective will be rotated according to the input video rotation rotation, this configuration is per instance.Image source rotation

Object Snapshot feature​

A complete list of devices that supports this feature can be found at the AXIS Scene Metadata product page by filtering on the functionality "Object Snapshot".

The Object Snapshot feature enables the capturing of object snapshots in a scene. When enabled this feature adds a base-64 encoded cropped image to both classified and unclassified objects. The current implementation is heavily focused on bounding box size during the track lifetime. Snapshots may be sent more than once for each object, if better alternatives are found.

This feature interfaces with the Fusion Tracker module as shown below,

Compression​

This feature uses resolution downscaling and JPEG quality compression to fit the snapshot generation requirements. The JPEG quality does not go below 50% to ensure a useful snapshot.

To enable this feature see guide.

Snapshot generation requirements​

All the requirements are checked after the image compression step

  • Size less than 64 kB
  • Pixel count
    • Less than 500k
    • More than 800

Snapshot limitations​

  • Tracks consisting of less than four frames are unlikely to receive a snapshot

Radar video fusion feature​

info

This feature is only available radar video fusion cameras such as Q1686-DLE and Q1656-DLE.

The radar video fusion feature enables the Fusion Tracker module to also fuse tracking data generated from radar scans.

This feature will add additional information to the object state for objects detected by the radar. The additional info includes

  • Spherical Coordinate
  • Speed
  • Direction

As the radar sensor is designed to detect objects far away, using this feature will also improve detection range.

This feature can also be used in combination with the Geographic Coordinates this feature to receive geolocation data.

License plate verifier integration feature​

info

This feature is only available on Q1686-DLE.

The AXIS License Plate Verifier ACAP enables capturing and recognizing of license plate text. By integrating this data into the Fusion Tracker module through this feature, the same information can now be provided for the Fusion Tracker module output.

The ALPV ACAP acap can be visualized as a processing module together with the Fusion Tracker module as such,

Geographic coordinates​

info

This feature is only available radar video fusion cameras such as Q1686-DLE and Q1656-DLE.

The geographic coordinates feature adds coordinates in decimal degrees to the object state. To calculate accurate coordinates, configure the device's location and orientation precisely. The bottom center of each object's bounding box is used to calculate its coordinates. If a person's feet aren't visible, the coordinates may be inaccurate.

This feature uses data from the Radar Motion Tracker. Only objects detected by the radar in a given frame receive geolocation data for that frame.

Read the how-to on how to enable this feature.

Limitations​

  • Mirroring is not supported.
    • If the object tracking should be visualized on a mirrored image, the user must apply mirroring.
  • Objects must be moving to be considered by this tracker.
  • This tracker only supports video stream default aspect ratios.

Consuming data from multiple channels​

On devices with multiple video sensors, or fisheye cameras that support dewarped views, consuming scene metadata from three or more channels simultaneously may result in dropped messages. To reduce this risk, consume data from no more than two channels at a time.