DRONE EMOTIONS RESEARCH · BENCHMARK PROTOCOL 006

Large Dataset Performance Benchmarking in Agisoft Metashape

A reproducible protocol for comparing Metashape processing performance across hardware, workflow settings, software versions and deployment architectures while preserving dataset integrity, output quality and transparent test conditions.

Review the Benchmark Protocol Discuss a Benchmark

06
BENCHMARK PROTOCOL

7 BENCHMARK STAGES

Define · Freeze · Prepare · Run
Monitor · Validate · Report

DEVELOPED BY DRONE EMOTIONS

Last reviewed: September 2026

TRANSPARENT CONTENT CLASSIFICATION

This is an independent benchmarking methodology developed by Drone Emotions Srl. It does not present measured benchmark results, guaranteed processing times or a universal ranking of hardware. Valid results require a frozen dataset, controlled workflow, documented environment, repeated observations and verification that compared outputs remain technically equivalent.

THE CORE PRINCIPLE

A faster result is meaningful only when the comparison is fair

Processing time alone cannot establish that one system or workflow is better. A setting that finishes more quickly may produce a different point density, model detail, resolution, texture, accuracy or level of completeness.

A defensible benchmark changes one defined factor, holds the remaining conditions as constant as practical, records resource behaviour by stage and confirms that the compared outputs still satisfy the same acceptance criteria.

✓ Start with one explicit benchmark question

✓ Freeze the dataset, workflow and acceptance criteria

✓ Measure resources and stage duration—not only total time

✓ Validate output equivalence before ranking results

BENCHMARK QUESTIONS

What can be compared?

Hardware
CPU, GPU, RAM, storage or complete systems.

Workflow settings
Quality, filtering, preselection or output parameters.

Software environment
Metashape versions, drivers or operating systems.

Architecture
Workstation, network processing, server, cloud or hybrid deployment.

FOUR CONTROL DOMAINS

Control the dataset, environment, workflow and result

If one domain changes without being recorded, the comparison may no longer support the intended conclusion.

01

Dataset

Exact source files, metadata, reference data, masks, chunk structure and integrity.

02

Environment

Software build, OS, drivers, hardware, power profile, storage path and background load.

03

Workflow

Stage order, parameters, calibration, region, intermediate reuse and automation.

04

Outcome

Duration, resource peaks, failures, output statistics, accuracy, completeness and quality.

BENCHMARK SEQUENCE

Seven stages from research question to published evidence

The sequence is designed to make later repetition and comparison possible.

01

Define

State the variable, hypothesis, metrics and acceptance rule.

02

Freeze

Lock source data, project baseline and workflow parameters.

03

Prepare

Document and stabilize the complete test environment.

04

Run

Execute the same controlled stages and retain logs.

05

Monitor

Capture duration, utilization, resource peaks, temperatures and failures by stage.

06

Validate

Confirm output completeness, quality and accuracy remain comparable.

07

Report

Publish conditions, repeated results, variability, evidence and limitations.

DETAILED BENCHMARK PROTOCOL

Ten controls for a reproducible comparison

Document every condition that could materially affect the result. If a comparison cannot be reproduced or its outputs are not equivalent, it should not be presented as a definitive benchmark.

01 — Define one research question and comparison variable

A benchmark should answer a precise question rather than generate an isolated time.

  • State whether the test compares hardware, settings, software versions or architectures.
  • Identify the single primary variable intended to change.
  • Define primary metrics such as stage duration, peak RAM, VRAM, storage growth or energy where available.
  • Define output-quality and accuracy criteria that must remain satisfied.
  • State the hypothesis before the run.
  • Define repetition count and treatment of failed or interrupted runs.
  • Identify the audience and the conclusions the test is not designed to support.
  • Register the test identifier, author, reviewer and date.

Benchmark record: research question, controlled variable, metrics, acceptance criteria and exclusions.

The dataset determines how widely the result can be interpreted.

  • Record image count, dimensions, megapixels, bit depth, format and total volume.
  • Describe aerial, corridor, oblique, terrestrial, close-range or mixed capture geometry.
  • Record cameras, lenses, sensors, flights, sessions and metadata.
  • Describe scene complexity, texture, vegetation, reflections, movement and elevation range.
  • List GCPs, checkpoints, masks, laser scans or other inputs.
  • State why the dataset represents the intended production workload.
  • Separate “large by image count” from “large by resolution, geometry or output complexity”.
  • Confirm permission to use and, if applicable, publish information about the dataset.

Benchmark record: dataset profile, permitted use and explanation of representativeness.

Every compared run must begin from the same controlled state.

  • Create a read-only or otherwise controlled copy of source files.
  • Record file counts, folder structure and integrity checks where required.
  • Preserve metadata and reference files without conversion between runs.
  • Freeze masks, calibration groups, camera status and marker projections.
  • Define the project state from which each stage begins.
  • State whether intermediate data are rebuilt or reused.
  • Clear or control cached results when they could affect timing.
  • Use consistent region, chunks and enabled items.

Benchmark record: immutable input manifest and restorable project baseline.

A component name alone is not a complete test environment.

  • Record Metashape edition, version and build.
  • Record operating system, updates and power profile.
  • Record CPU model, core configuration, memory capacity and memory configuration.
  • Record GPU model, VRAM, enabled devices and driver version.
  • Record active project storage, filesystem, free space and connection type.
  • For network or cloud tests, record server, workers, shared storage, network and instance details.
  • Record relevant firmware, BIOS and thermal-management settings.
  • Identify monitoring tools and their sampling interval.

Benchmark record: environment manifest sufficient to understand and repeat the setup.

The settings actually executed—not the intended settings—must match the record.

  • Record the exact order of processing stages.
  • Record matching, alignment, optimization, depth-map and filtering parameters.
  • Record point-cloud, mesh, DEM, orthomosaic, tiled-model and texture settings as applicable.
  • Record camera calibration, reference accuracy and control/checkpoint state.
  • Use the same batch configuration or script where practical.
  • Record whether CPU and GPU devices are enabled for supported stages.
  • Prevent unrecorded manual edits between comparable runs.
  • Export or capture the parameter record with the result package.

Benchmark record: executable workflow specification and verified run parameters.

Background load, thermal state and storage conditions can distort the comparison.

  • Close or record unrelated applications and scheduled tasks.
  • Use a consistent power and performance profile.
  • Record ambient and component temperatures where they may affect throttling.
  • Define whether runs begin from a cold, warm or stabilized system state.
  • Maintain comparable free storage and project location.
  • Control network traffic and concurrent storage use for distributed tests.
  • Record reboots, driver resets, failed tasks, retries and operator intervention.
  • Avoid changing hardware configuration between repetitions unless it is the tested variable.

Benchmark record: pre-run checklist and execution-condition log.

Total project time can hide where a difference actually occurs.

  • Record start, finish and elapsed time for every tested stage.
  • Capture peak and representative system RAM use.
  • Capture CPU utilization, frequency and temperature behaviour.
  • Capture GPU utilization, VRAM allocation, clock and temperature behaviour.
  • Record storage read/write activity, project growth and minimum free space.
  • Record network throughput and worker utilization where applicable.
  • Separate automated processing time from manual review or transfer time.
  • Retain Metashape and system logs required to explain anomalies.

Benchmark record: stage-level telemetry and event log synchronized with each run.

One successful run may not represent typical behaviour.

  • Perform the predefined number of comparable repetitions.
  • Use the same reset or preparation procedure before each run.
  • Report individual observations rather than only the best result.
  • Calculate an appropriate summary and variability measure for the test design.
  • Investigate outliers before deciding whether they are valid, failed or excluded.
  • Record thermal drift, cache effects, background events and storage changes.
  • Repeat both comparison conditions under equivalent rules.
  • State when resource cost or duration prevented sufficient repetition.

Benchmark record: all repeated observations, variability and documented treatment of anomalies.

A performance gain is not comparable if the result is materially different.

  • Compare aligned camera count, tie points and calibration state.
  • Compare checkpoint or scale-bar evidence where relevant.
  • Compare depth-map, point-cloud, mesh, DEM, orthomosaic or texture statistics.
  • Inspect completeness, holes, noise, artefacts and spatial coverage.
  • Confirm resolution, region, CRS and export settings are equivalent.
  • Use identical acceptance criteria for all compared runs.
  • Identify nondeterministic differences and assess whether they affect the conclusion.
  • Reject a ranking when output equivalence cannot be demonstrated.

Benchmark record: output-equivalence assessment linked to every accepted run.

A reader should be able to understand exactly what the benchmark proves.

  • Publish the research question and comparison variable.
  • Describe the dataset without exposing restricted information.
  • List software, hardware, drivers, settings and execution conditions.
  • Report every valid repetition and summary method.
  • Report failures, exclusions, anomalies and output-equivalence checks.
  • Separate observed results from interpretation and recommendation.
  • State that results apply to the tested conditions and may not scale linearly.
  • Provide a review date and preserve the benchmark package.

Benchmark record: transparent report plus archived dataset manifest, workflow, logs and validation evidence.

FOUR BENCHMARK DESIGNS

Match the test structure to the decision

Each design requires different controls. Combining several questions in one run makes causal interpretation weaker.

DESIGN A

Hardware comparison

Same data, software and workflow across different components or complete systems.

Control: drivers, power, storage and thermal conditions.

DESIGN B

Settings comparison

Same environment and data with one workflow parameter deliberately changed.

Control: output quality and acceptance equivalence.

DESIGN C

Version comparison

Same data and environment processed with defined Metashape versions.

Control: compatibility, defaults, algorithms and project state.

DESIGN D

Architecture comparison

Workstation, network, server or cloud execution under documented conditions.

Control: transfer, shared storage, workers and total operational time.

BENCHMARK INPUT RECORD

Conditions that must be frozen

Dataset
Files, integrity, metadata, reference, masks, groups and project state.

Workflow
Stage order, parameters, region, calibration and reuse rules.

Environment
Software, OS, drivers, hardware, storage, network and power.

Execution
Background load, temperature state, repetitions and monitoring method.

BENCHMARK RESULT RECORD

Evidence that must be measured

Performance
Stage duration, total active time, retries and operator intervention.

Resources
RAM, CPU, GPU, VRAM, storage, network and temperature behaviour.

Reliability
Warnings, errors, failed operations, throttling and recovery.

Quality
Output statistics, accuracy, completeness, artefacts and acceptance.

COMPARISON STATUS

Comparable, conditional or invalid

Classify the comparison before interpreting which result appears faster.

COMPARABLE

Conditions and outputs are equivalent

The intended variable changed, repetitions are valid and all results satisfy the same technical criteria.

CONDITIONAL

The comparison has disclosed constraints

Minor environmental or output differences limit the conclusion but remain sufficiently documented.

INVALID

The result cannot support a ranking

Uncontrolled variables, failed repetitions, missing telemetry or non-equivalent outputs invalidate the comparison.

LIMITATIONS AND RESPONSIBLE USE

Benchmark results belong to their test conditions

Performance and resource demand depend on dataset content, image resolution, processing parameters, output type, software version, drivers, hardware, storage, system state and deployment architecture. Results from one dataset cannot be assumed to scale linearly to another.

Benchmark data should not be presented as a purchasing guarantee. Production capacity requires representative testing, operational headroom, reliability assessment and consideration of support, licensing, backup and security.

✓ No performance result is fabricated

✓ No processing time is guaranteed

✓ No hardware ranking is presented

✓ Output equivalence is mandatory

Agisoft and Metashape are trademarks of their respective owner. This independent resource is developed by Drone Emotions Srl, an Agisoft Authorized Reseller and Training Center, and is not official Agisoft LLC documentation.

PRIMARY TECHNICAL REFERENCES

Compare your method with current Agisoft guidance

Official tests and recommendations provide useful reference conditions, but their results must be interpreted within the documented dataset and environment.

AGISOFT HELPDESK

Memory Requirements

Official benchmark conditions and memory observations by processing stage.

Open Official Guide
AGISOFT HELPDESK

Hardware Recommendations

Current guidance for storage, RAM, CPU, GPU and stability.

Open Official Guide
AGISOFT HELPDESK

GPU Processing

Official description of GPU-supported operations and configuration.

Open Official Guide
AGISOFT HELPDESK

Network Processing

Official client, server, worker and shared-storage configuration.

Open Official Guide
TECHNICAL GUIDE 003

Hardware Planning

Profile workloads and plan RAM, GPU, CPU, storage and deployment before procurement.

Open the Guide
QA PROTOCOL 004

Reproducible Workflows

Preserve source provenance, processing states, quality gates, reports and releases.

Review the Protocol
DIAGNOSTIC PROTOCOL 005

Camera Diagnostics

Review image quality, calibration groups, parameter behaviour and independent evidence.

Review the Protocol

QUESTIONS

About performance benchmarks

A benchmark is evidence about a defined test—not a universal promise.

No. Dataset content, image resolution, geometry, settings, outputs, software and infrastructure can all change performance. Results apply to the documented test conditions.
Repeated runs reveal variability caused by temperature, cache, background activity, storage and other system conditions. One run may be unusually fast or slow.
Not as a simple performance ranking. The quality difference must be reported, and the outputs must satisfy the same acceptance criteria before speed alone can be compared fairly.
No. Stage time, peak memory, GPU and CPU behaviour, storage, reliability, failures and final output quality help explain why a result changed.
Sometimes, but only if it preserves the characteristics that create the workload. Scaling may be nonlinear, so the limits of extrapolation must be stated.
No. It defines the method Drone Emotions can use for future measured research. Results should be published only after controlled tests have been completed and reviewed.
No. It is an independent protocol developed by Drone Emotions Srl. Official Agisoft tests and documentation should be consulted as primary product references.

DRONE EMOTIONS RESEARCH

Turn performance claims into reproducible evidence

Describe the dataset, systems or settings you need to compare. Drone Emotions can help define a controlled benchmark question, measurement plan and transparent reporting structure.

Email the Technical Team