Synthetic data generation and augmentation for biomechanical time-series data.

Christina Halmich (2026): Synthetic data generation and augmentation for biomechanical time-series data. Methods, Challenges and Perspectives. In: dvs-Spinfortec 2026 Tagungsband: Beiträge zur Tagung der dvs-Sektion Sportinformatik und Sporttechnologie 2026 an der Otto-von-Guericke-Universität Magdeburg (S. 86–87). Steinbeis-Edition. ISBN 978-3-95663-337-9

The integration of machine learning and deep learning into human movement biomechanics has gained increased relevance in recent years. Application areas range from joint angle estimation to fall detection and movement analysis in sport and rehabilitation. However, these models require large, diverse, and well-annotated training data sets to perform reliably on unseen data (Knudson, 2017). Acquiring biomechanical datasets is resource-intensive and often limited by laboratory requirements, participant availability, ethical and data privacy constraints, sensor-related issues, and restricted movement variability (Knudson, 2017; Halmich et al., 2025). These obstacles can result in models that perform reliably only under narrowly defined conditions and show reduced generalizability across individuals and tasks (Sharifi Renani et al., 2021).

Data augmentation has emerged as a promising approach to address this scarcity (Wen et al., 2021). By synthetically generating or systematically modifying training data, data augmentation has been shown to improve model performance and robustness in many studies (Halmich et al., 2025). Yet, compared to image-based domains, where augmentation pipelines are well established, equivalent methodology for biomechanical time-series data remains less standardized and insufficiently evaluated (Wen et al., 2021).

Drawing on findings from our scoping review (Halmich et al., 2025), this contribution aims to raise the topic of synthetic data augmentation for biomechanical time series data and stimulate discussion on its methodological challenges and open questions.

Methods and Challenges

Based on existing literature (Halmich et al., 2025), augmentation methods for biomechanical time-series data can be grouped into three main categories, physics-based methods (e.g. musculoskeletal models), classical methods (e.g., jittering), and data-driven methods (e.g. Generative Adversarial Networks).

Despite this variety of approaches, systematic comparisons and evaluations of augmentation methods remain limited, making it difficult to assess their respective strengths and weaknesses (Halmich et al., 2025). The majority of studies assess methods based on the accuracy of downstream models trained on augmented vs. non-augmented data for a specific task (e.g. gait phase classification). Direct comparison with measured data and biomechanical validity assessment are often omitted. This raises the concern that methods generating primarily statistical noise may improve downstream performance in the short term, without the generated data being biomechanically meaningful (Halmich et al., 2025).

A further challenge is the direct transfer of augmentation strategies from image processing to biomechanical time-series data (Wen et al., 2021). Generative modelling approaches are often adopted from domains in which transformations can be assessed more intuitively, particularly image-based applications (Wen et al., 2021). IMU signals, however, are shaped by biomechanical and contextual factors (e.g. sensor position and orientation), environmental conditions (e.g. incline, movement speed) and individual characteristics (e.g. body mass, height). An augmentation method that does not account for these dependencies risks generating biomechanically implausible signals.

Another challenge concerns the training of data-driven generative methods. While such methods are intended to address data scarcity by generating additional samples, they themselves typically require sufficiently large and representative datasets to learn meaningful temporal dependencies and biomechanical constraints (Li et al., 2022). This creates a methodological dilemma: the models are used because data are limited, but their training may be constrained by the same lack of data they are meant to overcome.

Scientific Relevance and Outlook

Whether in running, skiing, or clinical rehabilitation, comparable constraints apply: small participant numbers, high sensor sensitivity to placement and movement, as well as limited variability in movement. Across all of these domains, practitioners face the same fundamental bottleneck – insufficient data to train models that generalize well (Knudson, 2017).

Even though methods from the image domain cannot be directly transferred, there are a few approaches one can adapt for time-series data. Methods such as conditioned data augmentation (Li et al., 2022) and curriculum learning may provide useful starting points. These methods may help us address data scarcity by producing realistic samples, improve transferability by generating data across a wider range of conditions and produce meaningful samples that reflect genuine biomechanical structure rather than statistical noise.

Addressing these challenges requires moving augmentation research in biomechanics from ad hoc application toward methods that are evaluated, and understood, on biomechanical terms. The path forward lies not in importing methods from other domains, but in developing augmentation approaches that are accountable to biomechanical structure from the outset.

Publikationsautor:innen der Salzburg Research (in alphabetischer Reihenfolge):

Link

Relevante Projekte:

Newsletter
Erhalten Sie dreimal jährlich unseren postalischen Newsletter sowie Einladungen zu Veranstaltungen. Kostenlos abonnieren.

Kontakt
Salzburg Research Forschungsgesellschaft
Jakob Haringer Straße 5/3
5020 Salzburg, Austria
Top