References#
Every paper, dataset and tool cited in these docs. A citation on any page links to its entry here.
Software EthoGraph builds on#
Jan Benda. Audioio: platform independent interfacing of numpy arrays of floats with audio files and devices. URL: bendalab/audioio.
Samuel Garcia, Domenico Guarino, Florent Jaillet, Todd Jennings, Robert Pröpper, Philipp L. Rautenberg, Chris C. Rodgers, Andrey Sobolev, Thomas Wachtler, Pierre Yger, and Andrew P. Davison. Neo: an object model for handling electrophysiology data in multiple formats. Frontiers in Neuroinformatics, 8:10, 2014. doi:10.3389/fninf.2014.00010.
Stephan Hoyer and Joe Hamman. Xarray: N-D labeled arrays and datasets in Python. Journal of Open Research Software, 5(1):10, 2017. doi:10.5334/jors.148.
Almar Klein and others. Pygfx: a render engine for Python. URL: https://pygfx.org/, doi:10.5281/zenodo.21443465.
David Nicholson. Crowsetta: a Python tool to work with any format for annotating animal vocalizations and bioacoustics data. Journal of Open Source Software, 8(84):5338, 2023. doi:10.21105/joss.05338.
Cyrille Rossant, Shabnam N. Kadir, Dan F. M. Goodman, John Schulman, Maximilian L. D. Hunter, Aman B. Saleem, Andres Grosmark, Mariano Belluscio, George H. Denfield, Alexander S. Ecker, Andreas S. Tolias, Samuel Solomon, György Buzsáki, Matteo Carandini, and Kenneth D. Harris. Spike sorting for large, dense electrode arrays. Nature Neuroscience, 19:634–641, 2016. The paper behind phy (https://github.com/cortex-lab/phy). doi:10.1038/nn.4268.
Oliver Rübel, Andrew Tritt, Ryan Ly, Benjamin K. Dichter, Satrajit Ghosh, Lawrence Niu, Pamela Baker, Ivan Soltesz, Lydia Ng, Karel Svoboda, Loren Frank, and Kristofer E. Bouchard. The Neurodata Without Borders ecosystem for neurophysiological data science. eLife, 11:e78362, 2022. doi:10.7554/eLife.78362.
Nikoloz Sirmpilatze, Chang Huan Lo, Sofía Miñano, Brandon D. Peri, Dhruv Sharma, Laura Porta, Iván Varela, and Adam L. Tyson. Movement. URL: https://movement.neuroinformatics.dev/, doi:10.5281/zenodo.12755724.
Guillaume Viejo, Daniel Levenstein, Sofia Skromne Carrasco, Dhruv Mehrotra, Sara Mahallati, Gilberto R. Vite, Henry Denny, Lucas Sjulson, Francesco P. Battaglia, and Adrien Peyrache. Pynapple, a toolbox for data analysis in neuroscience. eLife, 12:RP85786, 2023. doi:10.7554/eLife.85786.
PyAV contributors. PyAV: pythonic bindings for FFmpeg's libraries. URL: https://pyav.org/docs/stable/.
pynapple-org. Pynaviz: interactive visualization for pynapple. URL: pynapple-org/pynaviz.
PyQtGraph contributors. PyQtGraph: scientific graphics and GUI library for Python. URL: https://www.pyqtgraph.org/.
Datasets#
Jesse Marshall, Ugne Klibaite, Amanda Gellis, Diego Aldarondo, Bence Olveczky, and Timothy W. Dunn. The PAIR-R24M dataset for multi-animal 3D pose estimation. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1. 2021. URL: https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/1ff8a7b5dc7a7d1f0ed65aaa29c04b1e-Abstract-round1.html.
Felix W. Moll, Julius Würzler, and Andreas Nieder. Learned precision tool use in carrion crows. Current Biology, 35(19):4845–4852.e3, 2025. doi:10.1016/j.cub.2025.08.033.
Tanmay Nath, Alexander Mathis, An Chi Chen, Amir Patel, Matthias Bethge, and Mackenzie Weygandt Mathis. Using DeepLabCut for 3D markerless pose estimation across species and behaviors. Nature Protocols, 14(7):2152–2176, 2019. doi:10.1038/s41596-019-0176-0.
Patrik Reiske, Marcus N. Boon, Niek Andresen, Sole Traverso, Marieatou Daniels, Katharina Hohlbaum, Lars Lewejohann, Christa Thöne-Reineke, Olaf Hellwich, and Henning Sprekeler. Mouse lockbox dataset: behavior recognition for mice solving lockboxes. International Journal of Computer Vision, 134(7):318, 2026. doi:10.1007/s11263-026-02908-x.
Linus Rüttimann, Yuhang Wang, Jörg Rychen, Tomas Tomka, Heiko Hörster, and Richard H. R. Hahnloser. Multimodal system for recording individual-level behaviors in songbird groups. PeerJ, 13:e20203, 2025. doi:10.7717/peerj.20203.
Changepoint detection#
Yarden Cohen, David Aaron Nicholson, Alexa Sanchioni, Emily K. Mallaber, Viktoriya Skidanova, and Timothy J. Gardner. Automated annotation of birdsong with a neural network that segments spectrograms. eLife, 11:e63853, 2022. doi:10.7554/eLife.63853.
Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, and Richard H. R. Hahnloser. Positive transfer of the Whisper speech transformer to human and animal voice activity detection. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 7505–7509. 2024. doi:10.1109/ICASSP48485.2024.10447620.
David E. Meyer, Richard A. Abrams, Sheldon Kornblum, Charles E. Wright, and J. E. Keith Smith. Optimality in human motor performance: ideal control of rapid aimed movements. Psychological Review, 95(3):340–370, 1988. doi:10.1037/0033-295X.95.3.340.
David Nicholson. Vocalpy/vocalpy: 0.2.0. 2023. doi:10.5281/zenodo.7905426.
David Nicholson and Yarden Cohen. Vak: a neural network framework for researchers studying animal acoustic communication. In Proceedings of the 22nd Python in Science Conference (SciPy). 2023. doi:10.25080/gerudo-f2bc6f59-008.
Tim Sainburg, Marvin Thielk, and Timothy Q. Gentner. Finding, visualizing, and quantifying latent structure across diverse animal vocal repertoires. PLOS Computational Biology, 16(10):e1008228, 2020. doi:10.1371/journal.pcbi.1008228.
Francesco Torricelli, Alice Tomassini, Giovanni Pezzulo, Thierry Pozzo, Luciano Fadiga, and Alessandro D'Ausilio. Motor invariants in action execution and perception. Physics of Life Reviews, 44:13–47, 2023. doi:10.1016/j.plrev.2022.11.003.
Charles Truong, Laurent Oudre, and Nicolas Vayatis. Selective review of offline change point detection methods. Signal Processing, 167:107299, 2020. doi:10.1016/j.sigpro.2019.107299.
Ruiyu Xu, Zheren Song, Jianguo Wu, Chao Wang, and Shiyu Zhou. Change-point detection with deep learning: a review. Frontiers of Engineering Management, 12(1):154–176, 2025. doi:10.1007/s42524-025-4109-z.
Keypoint labelling and tracking#
Jean-Yves Bouguet. Pyramidal implementation of the lucas kanade feature tracker. 2001. URL: https://robots.stanford.edu/cs223b04/algo_tracking.pdf.
F. N. Fritsch and R. E. Carlson. Monotone piecewise cubic interpolation. SIAM Journal on Numerical Analysis, 17(2):238–246, 1980. doi:10.1137/0717021.
Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. CoTracker3: simpler and better point tracking by pseudo-labelling real videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 6013–6022. 2025. URL: https://arxiv.org/abs/2410.11831, arXiv:2410.11831.
Maximilian Krogius, Acshi Haggenmiller, and Edwin Olson. Flexible layouts for fiducial tags. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1898–1903. 2019. doi:10.1109/IROS40897.2019.8967787.
Bruce D. Lucas and Takeo Kanade. An iterative image registration technique with an application to stereo vision. In Proceedings of the 7th International Joint Conference on Artificial Intelligence (IJCAI), 674–679. 1981. URL: https://www.ri.cmu.edu/pub_files/pub3/lucas_bruce_d_1981_1/lucas_bruce_d_1981_1.pdf.
Zhuoyang Pan, Boxiao Pan, Guandao Yang, Adam W. Harley, and Leonidas Guibas. Animal pose labeling using general-purpose point trackers. In CV4Animals Workshop at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2025. URL: https://arxiv.org/abs/2506.03868, arXiv:2506.03868.
Point-event models#
James Hong, Haotian Zhang, Michaël Gharbi, Matthew Fisher, and Kayvon Fatahalian. Spotting temporally precise, fine-grained events in video. In European Conference on Computer Vision (ECCV). 2022. URL: https://arxiv.org/abs/2207.10213, arXiv:2207.10213.
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. LightGBM: a highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (NeurIPS). 2017. URL: https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree.
Zhaoyu Liu, Kan Jiang, Murong Ma, Zhe Hou, Yun Lin, and Jin Song Dong. Few-shot precise event spotting via unified multi-entity graph and distillation. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2026. URL: https://arxiv.org/abs/2511.14186, arXiv:2511.14186.
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake VanderPlas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Edouard Duchesnay. Scikit-learn: machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. URL: https://jmlr.org/papers/v12/pedregosa11a.html.
Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020. URL: https://arxiv.org/abs/2003.13678, arXiv:2003.13678.
Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz. Gate-shift networks for video action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020. URL: https://arxiv.org/abs/1912.00381, arXiv:1912.00381.
Hao Xu, Xinyu Wei, Sam Wells, and Sunil Aryal. MSAGSM: multi-scale attention gated shifting for precise event spotting. 2025. arXiv preprint; later versions retitled "Multi-Focus Temporal Shifting for Precise Event Spotting in Sports Videos". URL: https://arxiv.org/abs/2507.07381, arXiv:2507.07381.
Action segmentation#
Yazan Abu Farha and Juergen Gall. MS-TCN: multi-stage temporal convolutional network for action segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019. URL: https://arxiv.org/abs/1903.01945, arXiv:1903.01945.
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. Optuna: a next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD). 2019. URL: https://arxiv.org/abs/1907.10902, arXiv:1907.10902.
Emad Bahrami, Gianpiero Francesca, and Juergen Gall. How much temporal long-term context is needed for action segmentation? In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2023. URL: https://arxiv.org/abs/2308.11358, arXiv:2308.11358.
Haoyu Ji, Bowen Chen, Zhihao Yang, Wenze Huang, Yu Gao, Xueting Liu, Weihong Ren, Zhiyong Wang, and Honghai Liu. Spectral scalpel: amplifying adjacent action discrepancy via frequency-selective filtering for skeleton-based action segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2026. URL: HaoyuJi/SpecScalpel, arXiv:2603.24134.
Haoyu Ji, Xueting Liu, Yu Gao, Wenze Huang, Zhihao Yang, Weihong Ren, Zhiyong Wang, and Honghai Liu. LaDy: lagrangian-dynamic informed network for skeleton-based action segmentation via spatial-temporal modulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2026. URL: HaoyuJi/LaDy, arXiv:2603.24097.
Elizaveta Kozlova, Andy Bonnetto, and Alexander Mathis. DLC2Action: a multimodal deep learning-based toolbox for automated behavior segmentation. 2025. bioRxiv preprint. URL: amathislab/DLC2Action, doi:10.1101/2025.09.27.678941.
Colin Lea, Michael D. Flynn, René Vidal, Austin Reiter, and Gregory D. Hager. Temporal convolutional networks for action segmentation and detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017. URL: https://arxiv.org/abs/1611.05267, arXiv:1611.05267.
Shijie Li, Yazan Abu Farha, Yun Liu, Ming-Ming Cheng, and Juergen Gall. MS-TCN++: multi-stage temporal convolutional network for action segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. URL: https://arxiv.org/abs/2006.09220, arXiv:2006.09220.
Daochang Liu, Qiyue Li, AnhDung Dinh, Tingting Jiang, Mubarak Shah, and Chang Xu. Diffusion action segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2023. URL: https://arxiv.org/abs/2303.17959, arXiv:2303.17959.
Zijia Lu and Ehsan Elhamifar. FACT: frame-action cross-attention temporal modeling for efficient action segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18175–18185. 2024. URL: https://openaccess.thecvf.com/content/CVPR2024/html/Lu_FACT_Frame-Action_Cross-Attention_Temporal_Modeling_for_Efficient_Action_Segmentation_CVPR_2024_paper.html.
Dipika Singhania, Rahul Rahaman, and Angela Yao. Coarse to fine multi-resolution temporal convolutional network. 2021. URL: https://arxiv.org/abs/2105.10859, arXiv:2105.10859.
Fangqiu Yi, Hongyu Wen, and Tingting Jiang. ASFormer: transformer for action segmentation. In British Machine Vision Conference (BMVC). 2021. URL: https://arxiv.org/abs/2110.08568, arXiv:2110.08568.
Wentao Zhu, Xiaoxuan Ma, Zhaoyang Liu, Libin Liu, Wayne Wu, and Yizhou Wang. MotionBERT: a unified perspective on learning human motion representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2023. URL: https://arxiv.org/abs/2210.06551, arXiv:2210.06551.
Video features#
Vladimir Iashin. Video_features: extract features from videos with a single command. 2020. URL: v-iashin/video_features.
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and Andrew Zisserman. The kinetics human action video dataset. 2017. URL: https://arxiv.org/abs/1705.06950, arXiv:1705.06950.
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. DINOv2: learning robust visual features without supervision. Transactions on Machine Learning Research, 2024. URL: https://openreview.net/forum?id=a68SUt6zFt, arXiv:2304.07193.
Ross Wightman. Pytorch image models (timm). 2019. URL: huggingface/pytorch-image-models.
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: speed-accuracy trade-offs in video classification. In European Conference on Computer Vision (ECCV). 2018. URL: https://arxiv.org/abs/1712.04851, arXiv:1712.04851.