Effectively transferring this issue (since can't do it cross-org apparently) from the previously used private APL repo, posted on Sep 25, 2025
Some figures might not render due to permissions
From Neha:
Requesting to check for best practices in packaging the Sanes dataset (https://dandi.emberarchive.org/dandiset/000299).
Packaging script is https://github.com/aplbrain/BBQS-EMBER-Data-Ingest/blob/main/Sanes%20R34DA059513/sanesDatatoNWB_v2.ipynb
Cody response 1
Some notes on a first pass of this example session:
(1) Why were the video files duplicated and split into multiple parts? The contiguous (2 GB) video looks good as it is. Even if SLEAP was separately applied piecewise for some reason, these should likely be recombined into a single individual pose track with the appropriately combined timestamps. This will help significantly in reducing overall size of the dataset.
(2) The timestamps of the current PoseEstimation tracks are almost certainly not correct (should be time in seconds relative to global start time) since they are all integers with typical step size of 1 (sometimes more). Perhaps they are meant to be frame indices whose timestamps can then be 'inferred' through the timing information of the linked ImageSeries? However, the videos in question only have starting time and rate so this may not be super well synchronized with the audio.
(3) The video and audio streams have the same starting_time - is this true? No timestamps are set on any objects, implying all synchronization is done by extrapolating the timings from the approximate rates of each data stream rather than by TTLs on frame captures. Do you trust this in analysis? Especially given the high rate of the acoustic data (125 kHz) I would tend to expect significant drift over long sessions.
(4) The acoustic data should be annotated using the AcousticWaveformSeries of the ndx-sound extension, instead of a basic TimeSeries. This will allow external tools (NWB Widgets, possibly Neurosift one day) to recognize it as audio and allow automatic visualizations such as spectrograms and playback.
(5) The acoustic data is uncompressed. Please apply at least GZIP compression. All the pose estimation data (and timestamps) are also uncompressed. You can easily apply blanket compression using the default Backend Configuration tool through NeuroConv.
(6) I do apologize for this one; it is something we hope to fix in the coming months. I would directly write the output NWB files and video files in DANDI-compliant naming and layout and skip the dandi organize step. This will allow you to give the video files much more meaningful names, potentially even matching how they were originally. The only requirement is that each video needs a unique naming convention that will separate it from every other file on the archive (which is what dandi organize enforces in an extreme manner). Let me know if you need help with this.
(7) The labeled_videos and original_videos paths of the pose estimations are absolute paths and thus only true on the original file system used during conversion; please adjust these to be relative (and adjusted to whatever names you chose to match the DANDI layout).
(8) What is the purpose of the multiple 'tracks' per pose estimation? I have not seen that pattern before. I notice each series has slightly different data and sometimes different timestamps as well...
(9) This is the first file we've seen that pushes the boundaries on the multi-subject issue. I'd recommend using the layout we proposed for the multi-subject extension. This is 'kind of' done partially here since I do notice that special Subject object in the processing module - but the current example file only has the skeleton of a single subject. The subject metadata should then be of the classic type and be specific to just that individual, and the only data in this file should be the processed pose for that one subject (with the video and audio data going to the separate combined 'group NWB' file, which then contains no subject-specific pose data). This is the easiest way of currently assigning subject specificity per data type.
(10) The values of channel on the processed 'vocalizations' object seems to only have the values 0.0 and 1.0 - first, these should probably be integers. Second, should they not span all 12 microphones? I recommend adding a useful description to the column as well to explain this meaning further.
(11) The vocalizations table has a lovely Neurosift rendering (see below or from this link). However it does reveal that there is only one occurrence each of noise_proposals and vox_proposals - is this correct? If so, then disregard. It merely caught my eye as an oddity.
(12) Each of the pose estimation series, once adapted to the requested multi-subject format, should be nested further under the behavior submodule of processing. Thing of both of the higher level groups as 'tags' on the containers.
The video files themselves are all rigid field of view using H.264 codec in MP4 container, so that is great to see.
I have attached the full NWB Inspector report, which may highlight some additional issues or just repeat some of what I've said. In general I suggest trying to resolve as many issues as possible that are raised in its reports.
sub-Gerbil-001-002-003_inspection.txt
Side note: these gerbils are just too cute. All science aside, it's fun to just watch them move around and interact passively.
Neha response 1
> (1) Why were the video files duplicated and split into multiple parts? The contiguous (2 GB) video looks good as it is. Even if SLEAP was separately applied piecewise for some reason, these should likely be recombined into a single individual pose track with the appropriately combined timestamps. This will help significantly in reducing overall size of the dataset.
The original files were actually separate video files. It seems they ran SLEAP on each video separately, and generated the .slp. I used neurconv to convert those .slp into ndx-pose and added each one to the file. This is probably my own technical limitation, not sure what's the best way to concatenate all the data together. Maybe using sleap-io directly?
Example of original data structure:

(2) The timestamps of the current PoseEstimation tracks are almost certainly not correct (should be time in seconds relative to global start time) since they are all integers with typical step size of 1 (sometimes more). Perhaps they are meant to be frame indices whose timestamps can then be 'inferred' through the timing information of the linked ImageSeries? However, the videos in question only have starting time and rate so this may not be super well synchronized with the audio.
(3) The video and audio streams have the same starting_time - is this true? No timestamps are set on any objects, implying all synchronization is done by extrapolating the timings from the approximate rates of each data stream rather than by TTLS on frame captures. Do you trust this in analysis? Especially given the high rate of the acoustic data (125 kHz) I would tend to expect significant drift over long sessions.
I'll have to check the timestamps when fixing (1) as well. The videos and audio are synchronized together in each of the folders. when I concatenated the audio and video together, I'm assuming everything is synchronized.
(4) The acoustic data should be annotated using the AcousticWaveformSeries of the ndx-sound extension, instead of a basic TimeSeries. This will allow external tools (NWB Widgets, possibly Neurosift one day) to recognize it as audio and allow automatic visualizations such as spectrograms and playback.
I will set it as AcousticWaveFormSeries in the next upload attempt.
(6) I do apologize for this one; it is something we hope to fix in the coming months. I would directly write the output NWB files and video files in DANDI-compliant naming and layout and skip the dandi organize step. This will allow you to give the video files much more meaningful names, potentially even matching how they were originally. The only requirement is that each video needs a unique naming convention that will separate it from every other file on the archive (which is what dandi organize enforces in an extreme manner). Let me know if you need help with this.
I'll try that!
(7) The labeled_videos and original_videos paths of the pose estimations are absolute paths and thus only true on the original file system used during conversion; please adjust these to be relative (and adjusted to whatever names you chose to match the DANDI layout).
Will do.
(8) What is the purpose of the multiple 'tracks' per pose estimation? I have not seen that pattern before. I notice each series has slightly different data and sometimes different timestamps as well...
There is one track for each animal in the video. That's interesting about the timestamps, but may be related to the issue from (2) as well.
(9) This is the first file we've seen that pushes the boundaries on the multi-subject issue. I'd recommend using the layout we proposed for the multi-subject extension. This is 'kind of' done partially here since I do notice that special Subject object in the processing module - but the current example file only has the skeleton of a single subject. The subject metadata should then be of the classic type and be specific to just that individual, and the only data in this file should be the processed pose for that one subject (with the video and audio data going to the separate combined 'group NWB' file, which then contains no subject-specific pose data). This is the easiest way of currently assigning subject specificity per data type.
I guess this has only one skeleton because it is the same skeleton used for all three subjects. Do you recommend here then using the most current version of ndx-multisubjects to store the video/audio (grouped), even though it hasn't been finalized or integrated into core NWB? and then store the individual pose data in separate .nwb files?
(10) The values of channel on the processed 'vocalizations' object seems to only have the values 0.0 and 1.0 - first, these should probably be integers. Second, should they not span all 12 microphones? I recommend adding a useful description to the column as well to explain this meaning further.
(11) The vocalizations table has a lovely Neurosift rendering (see below or from this link). However it does reveal that there is only one occurrence each of noise_proposals and vox_proposals - is this correct? If so, then disregard. It merely caught my eye as an oddity.
I will ask the team about these instances. Thanks for your close attention to detail!
Cody response 2
I'll have to check the timestamps when fixing (1) as well. The videos and audio are synchronized together in each of the folders. when I concatenated the audio and video together, I'm assuming everything is synchronized.
Oh my, I am starting to question that even more now
Neurosift has a basic 'temporal alignment' feature (doesn't yet apply to the external videos, but does to all internal data streams)
Or perhaps they are 'aligned' for that period of time, the audio just doesn't span the entire pose tracking?
Skeleton misunderstanding
EDIT: Same 'Skeleton' template is re-used for all subjects
I guess this has only one skeleton because it is the same skeleton used for all three subjects.
That is quite odd - certainly not how I would have done it.
But you may be right. It seems one subject was singled out for a couple of nodes, and then the others were merely marked as bodies
Neurosift
But the edge structure shows fully connected structure
FWIW a more standard approach (and perhaps a good data re-use case, if the lab is done with this dataset and ready to open it to the world) would be to redo the pose estimation with a distinct skeleton per subject. My understanding is that DLC/SLEAP leverages skeleton information within their model training, and while it's possible to have unconnected nodes I assume that would have some affect on performance... Not to mention the skeleton here treats separate bodies of separate subjects as if they were one entity (a 'rat king'?) 🤷
By chance is there a publication of any kind to refer to for this dataset? Is it still being collected and analyzed?
There is one track for each animal in the video. That's interesting about the timestamps, but may be related to the issue from (2) as well.
And my confusion deepens... Since each track uses the same skeleton that isn't really indicating or communicating that aspect in the metadata itself
If you load it into SLEAP directly how does it behave? Does it look as expected with the 'tracked points' properly following individual subjects?
I guess it could be that SLEAP was run correctly but the output data structures (pre or post NWB conversion) got their skeletons scrambled up
Do you recommend here then using the most current version of ndx-multisubjects to store the video/audio (grouped), even though it hasn't been finalized or integrated into core NWB? and then store the individual pose data in separate .nwb files?
I would, personally. This sounds like a great first use case for multi-subject NWB contents
even though it hasn't been finalized
I would pair this conversion with the extension to act as a driving force to getting it done (and NWBEP process at least initialized). I believe there was only one issue that was holding it back; the placement of the subject in the general group due to some bugs. If needed I'd recommend scheduling some one-on-ones with Ryan to get that figured out
This is a pretty common way of pushing the development of extensions based on need. That's how most of the best extensions have happened and the way the automatic namespacing in NWB files works means that it's perfectly fine as long as it passes validation checks
or integrated into core NWB?
To date I don't believe any extension has been 🙃
and then store the individual pose data in separate .nwb files?
The philosophy is that any data streams (e.g., the tracks you speak of) that uniquely correspond to a subject ought to be in their own subject-specific NWB file with all the rich metadata specific to that subject. These files would not use the multi-subject extension and would appear just like any other NWB file seen on DANDI
Then for any data streams that are not specific to a single subject (such as the raw video and audio array), those would use the new extension to indicate a combination of the subjects.
Neha response 2
It does look like they have one skeleton and it looks like it tracks fine in SLEAP
Image
That misalignment with pose estimation data might be because the timestamps are wrong.
OK, I'll go ahead and start populating the multi-subjects format
Cody response 3
> It does look like they have one skeleton and it looks like it tracks fine in SLEAP
Aaahh... 🤦
I see now
The skeleton is indeed properly designed and overlayed with each subject, but it is the SAME node/edge structure for each subject and hence it is 're-used' for each subject
Apologies for the confusion
One small suggestion, maybe just rename it 'Skeleton' in that case to help indicate there are not more skeletons expected
And if possible, rename the track_[int] to subject_[int] to make that connection clear
Thank you so much for confirming with that visualization
Probably also wouldn't hurt to just mention that Skeleton template re-use in the Pose description to help get ahead of any future potential confusion
As for the last item
The original files were actually separate video files. It seems they ran SLEAP on each video separately, and generated the .slp. I used neurconv to convert those .slp into ndx-pose and added each one to the file. This is probably my own technical limitation, not sure what's the best way to concatenate all the data together. Maybe using sleap-io directly?
I will think on this a bit
If it were a truly segmented trialized structure that would be perfectly fine
Given that the videos seem perfectly contiguous it seems like such a shame to just combine them. One would get better visualizations (and easier re-use) with that in basically all situations
Neha response 3
Are you saying it would be better to keep the contiguous videos separate?
Cody response 4
> Are you saying it would be better to keep the contiguous videos separate?
Sorry - I'm saying it would be better to combine all split video contents (and derived data) into single objects since they do seem contiguous. The usual excuse for keeping them separate is if they are not contiguous (e.g., trialized)
For example, consider if they did a longer session than this one and followed the same splitting pattern. As a data re-user would you rather have dozens of individual videos (and hundreds over the dataset) or just a few? Would you rather have a single PoseEstimationSeries that spans the duration of the session, or dozens of ones that chain together in a particular way (which is only obvious by introspection of each series' name or timestamps)?
BTW Did the team mention why they split the files to feed them into SLEAP? This is a different group than the ones that mentioned the video splitting for lazy memory purposes right?
Neha response 4
Yes, I agree it would be easier to work with one concatenated file. They didn't mention why they split the files, I can ask though. Yes, it's a different group than the one Grace has been working with you on.
Effectively transferring this issue (since can't do it cross-org apparently) from the previously used private APL repo, posted on Sep 25, 2025
Some figures might not render due to permissions
From Neha:
Cody response 1
Some notes on a first pass of this example session:(1) Why were the video files duplicated and split into multiple parts? The contiguous (2 GB) video looks good as it is. Even if SLEAP was separately applied piecewise for some reason, these should likely be recombined into a single individual pose track with the appropriately combined timestamps. This will help significantly in reducing overall size of the dataset.
(2) The
timestampsof the currentPoseEstimationtracks are almost certainly not correct (should be time in seconds relative to global start time) since they are all integers with typical step size of 1 (sometimes more). Perhaps they are meant to be frame indices whose timestamps can then be 'inferred' through the timing information of the linkedImageSeries? However, the videos in question only have starting time and rate so this may not be super well synchronized with the audio.(3) The video and audio streams have the same
starting_time- is this true? No timestamps are set on any objects, implying all synchronization is done by extrapolating the timings from the approximate rates of each data stream rather than by TTLs on frame captures. Do you trust this in analysis? Especially given the high rate of the acoustic data (125 kHz) I would tend to expect significant drift over long sessions.(4) The acoustic data should be annotated using the
AcousticWaveformSeriesof the ndx-sound extension, instead of a basicTimeSeries. This will allow external tools (NWB Widgets, possibly Neurosift one day) to recognize it as audio and allow automatic visualizations such as spectrograms and playback.(5) The acoustic data is uncompressed. Please apply at least GZIP compression. All the pose estimation
data(andtimestamps) are also uncompressed. You can easily apply blanket compression using the default Backend Configuration tool through NeuroConv.(6) I do apologize for this one; it is something we hope to fix in the coming months. I would directly write the output NWB files and video files in DANDI-compliant naming and layout and skip the
dandi organizestep. This will allow you to give the video files much more meaningful names, potentially even matching how they were originally. The only requirement is that each video needs a unique naming convention that will separate it from every other file on the archive (which is whatdandi organizeenforces in an extreme manner). Let me know if you need help with this.(7) The
labeled_videosandoriginal_videospaths of the pose estimations are absolute paths and thus only true on the original file system used during conversion; please adjust these to be relative (and adjusted to whatever names you chose to match the DANDI layout).(8) What is the purpose of the multiple 'tracks' per pose estimation? I have not seen that pattern before. I notice each series has slightly different data and sometimes different timestamps as well...
(9) This is the first file we've seen that pushes the boundaries on the multi-subject issue. I'd recommend using the layout we proposed for the multi-subject extension. This is 'kind of' done partially here since I do notice that special
Subjectobject in the processing module - but the current example file only has the skeleton of a single subject. The subject metadata should then be of the classic type and be specific to just that individual, and the only data in this file should be the processed pose for that one subject (with the video and audio data going to the separate combined 'group NWB' file, which then contains no subject-specific pose data). This is the easiest way of currently assigning subject specificity per data type.(10) The values of
channelon the processed 'vocalizations' object seems to only have the values0.0and1.0- first, these should probably be integers. Second, should they not span all12microphones? I recommend adding a useful description to the column as well to explain this meaning further.(11) The vocalizations table has a lovely Neurosift rendering (see below or from this link). However it does reveal that there is only one occurrence each of
noise_proposalsandvox_proposals- is this correct? If so, then disregard. It merely caught my eye as an oddity.(12) Each of the pose estimation series, once adapted to the requested multi-subject format, should be nested further under the
behaviorsubmodule ofprocessing. Thing of both of the higher level groups as 'tags' on the containers.The video files themselves are all rigid field of view using H.264 codec in MP4 container, so that is great to see.
I have attached the full NWB Inspector report, which may highlight some additional issues or just repeat some of what I've said. In general I suggest trying to resolve as many issues as possible that are raised in its reports.
sub-Gerbil-001-002-003_inspection.txt
Side note: these gerbils are just too cute. All science aside, it's fun to just watch them move around and interact passively.
Neha response 1
> (1) Why were the video files duplicated and split into multiple parts? The contiguous (2 GB) video looks good as it is. Even if SLEAP was separately applied piecewise for some reason, these should likely be recombined into a single individual pose track with the appropriately combined timestamps. This will help significantly in reducing overall size of the dataset.The original files were actually separate video files. It seems they ran SLEAP on each video separately, and generated the .slp. I used neurconv to convert those .slp into ndx-pose and added each one to the file. This is probably my own technical limitation, not sure what's the best way to concatenate all the data together. Maybe using sleap-io directly?
Example of original data structure:

I'll have to check the timestamps when fixing (1) as well. The videos and audio are synchronized together in each of the folders. when I concatenated the audio and video together, I'm assuming everything is synchronized.
I will set it as AcousticWaveFormSeries in the next upload attempt.
I'll try that!
Will do.
There is one track for each animal in the video. That's interesting about the timestamps, but may be related to the issue from (2) as well.
I guess this has only one skeleton because it is the same skeleton used for all three subjects. Do you recommend here then using the most current version of ndx-multisubjects to store the video/audio (grouped), even though it hasn't been finalized or integrated into core NWB? and then store the individual pose data in separate .nwb files?
I will ask the team about these instances. Thanks for your close attention to detail!
Cody response 2
Oh my, I am starting to question that even more now
Neurosift has a basic 'temporal alignment' feature (doesn't yet apply to the external videos, but does to all internal data streams)
Or perhaps they are 'aligned' for that period of time, the audio just doesn't span the entire pose tracking?
Skeleton misunderstanding
EDIT: Same 'Skeleton' template is re-used for all subjects
That is quite odd - certainly not how I would have done it.
But you may be right. It seems one subject was singled out for a couple of nodes, and then the others were merely marked as bodies
Neurosift
But the edge structure shows fully connected structureFWIW a more standard approach (and perhaps a good data re-use case, if the lab is done with this dataset and ready to open it to the world) would be to redo the pose estimation with a distinct skeleton per subject. My understanding is that DLC/SLEAP leverages skeleton information within their model training, and while it's possible to have unconnected nodes I assume that would have some affect on performance... Not to mention the skeleton here treats separate bodies of separate subjects as if they were one entity (a 'rat king'?) 🤷By chance is there a publication of any kind to refer to for this dataset? Is it still being collected and analyzed?And my confusion deepens... Since each track uses the same skeleton that isn't really indicating or communicating that aspect in the metadata itself
If you load it into SLEAP directly how does it behave? Does it look as expected with the 'tracked points' properly following individual subjects?
I guess it could be that SLEAP was run correctly but the output data structures (pre or post NWB conversion) got their skeletons scrambled up
I would, personally. This sounds like a great first use case for multi-subject NWB contents
I would pair this conversion with the extension to act as a driving force to getting it done (and NWBEP process at least initialized). I believe there was only one issue that was holding it back; the placement of the subject in the
generalgroup due to some bugs. If needed I'd recommend scheduling some one-on-ones with Ryan to get that figured outThis is a pretty common way of pushing the development of extensions based on need. That's how most of the best extensions have happened and the way the automatic namespacing in NWB files works means that it's perfectly fine as long as it passes validation checks
To date I don't believe any extension has been 🙃
The philosophy is that any data streams (e.g., the tracks you speak of) that uniquely correspond to a subject ought to be in their own subject-specific NWB file with all the rich metadata specific to that subject. These files would not use the multi-subject extension and would appear just like any other NWB file seen on DANDI
Then for any data streams that are not specific to a single subject (such as the raw video and audio array), those would use the new extension to indicate a combination of the subjects.
Neha response 2
It does look like they have one skeleton and it looks like it tracks fine in SLEAPImage
That misalignment with pose estimation data might be because the timestamps are wrong.
OK, I'll go ahead and start populating the multi-subjects format
Cody response 3
> It does look like they have one skeleton and it looks like it tracks fine in SLEAPAaahh... 🤦
I see now
The skeleton is indeed properly designed and overlayed with each subject, but it is the SAME node/edge structure for each subject and hence it is 're-used' for each subject
Apologies for the confusion
One small suggestion, maybe just rename it 'Skeleton' in that case to help indicate there are not more skeletons expected
And if possible, rename the
track_[int]tosubject_[int]to make that connection clearThank you so much for confirming with that visualization
Probably also wouldn't hurt to just mention that Skeleton template re-use in the Pose description to help get ahead of any future potential confusion
As for the last item
I will think on this a bit
If it were a truly segmented trialized structure that would be perfectly fine
Given that the videos seem perfectly contiguous it seems like such a shame to just combine them. One would get better visualizations (and easier re-use) with that in basically all situations
Neha response 3
Are you saying it would be better to keep the contiguous videos separate?Cody response 4
> Are you saying it would be better to keep the contiguous videos separate?Sorry - I'm saying it would be better to combine all split video contents (and derived data) into single objects since they do seem contiguous. The usual excuse for keeping them separate is if they are not contiguous (e.g., trialized)
For example, consider if they did a longer session than this one and followed the same splitting pattern. As a data re-user would you rather have dozens of individual videos (and hundreds over the dataset) or just a few? Would you rather have a single
PoseEstimationSeriesthat spans the duration of the session, or dozens of ones that chain together in a particular way (which is only obvious by introspection of each series'nameortimestamps)?BTW Did the team mention why they split the files to feed them into SLEAP? This is a different group than the ones that mentioned the video splitting for lazy memory purposes right?
Neha response 4
Yes, I agree it would be easier to work with one concatenated file. They didn't mention why they split the files, I can ask though. Yes, it's a different group than the one Grace has been working with you on.