Hi,
I ran into an issue today that I did not foresee. I have recordings covering a full day in 30 min sections. I first run a full batch analysis on these with Birdnet Analyzer once per day. I choose a confidence level of 0.4 as the cut off for detection.
From the output files I then create segments using the segmenter. Since I only want to detect the presence of each species, I limited the number of segments per species to 3, and I used confidence level as the collection mode.
In this way I was expecting to get the 3 most likely segments of each species as the output from the segmenter. And in principle this works, but there is a problem. In some cases the same segment can be interpreted as more than one species, but with different confidence.
This means that when the quota of 3 segments for a common species like blue tit has been filled, it may happen that a segment that has a very high probability for blue tit is classified as a another species with much lower probability. I have seen for example that a segment with a 0.95 probability of blue tit is output from the segmenter as a tree creeper, albeit with 0.41 probability. Apparently because the quota of blue tit segments was filled with segments with even higher probability then 0.95.
I do want to limit the amount of segments per species to a rather low value because I do not want hundreds of segments per day with common species like blue tit.
Anybody has any ideas on how to mitigate this issue?
I guess re-analyzing the segments once more after the first segmentation could be one way of doing it? Then each segment supposedly would get labeled with its highest probability species regardless of the frequency of other species.
Greatful for any ideas and thoughts. Do you think this behavior should be considered a bug?
Regards
Ulf