[Heig] HEIG Obscore Note - ready for final review/ comments on the Note #part 1

BONNAREL FRANCOIS gmail francois.bonnarel at gmail.com
Tue Sep 1 19:33:11 CEST 2026


Hi Ian,
  I think when Mireille talked of 3 categories she was talking of
     1 ) response functions
     2 ) the new category of products such as pdf , draws, and regions 
that you initially called " advanced data products"
     3 ) the data products currently covered by the ivoa : 
dataproduct_type vocabulary extension of initial ObsCore vocabulary.
   So she was not exactly doing the same distinction that you are doing 
below because some of the products in 3 may indeed been produced by 
"inverse methods" or sophisticated analysis of raw data.
Actually, thinking more about it I realize that up to now data products 
in 3 ) were always containing measurements initiated by a messenger of 
the sky and not quantities derived from analysis of them such as 
"distances", metallicities or MCM simulations of any physical quantity, 
such as some of the pdf, draws and regions included in the current 
definition in the proposed endorsed note
   So our proposal is to still distinguish the 3 categories by using 
different vocabularies even if terms belonging to any of them could be  
used in the dataproduct_type column of ObscOre. This is consistent with 
the tables in the draft and could be reflected in the lain text by very 
little changes
   This is now more a semantics issue than ObsCore strictly speaking, 
because beside this we admit that the 3 categories of data could be 
exposed via ObscOre under some conditions.
Cheers
François

PS : by the way I congratulate you and Janet for your retirement. As 
Mireille is also retiring today and Catherine and I retired officially 
two years ago, this note definitely  exposes a high density of 
end-career people !
Le 27/08/2026 à 17:33, Dr. Ian N. Evans via heig a écrit :
> Hi Mireille,
>
> I think this is an area where we might disagree ;-)
>
> In my view, these “advanced” data product types can play roles in all 
> of the 3 categories (by which I assume you are meaning data 
> processing, data selection, and data analysis) and so segregation by 
> role is difficult.
>
> Historically we used to acquire raw images, spectra, etc. at the 
> telescope and there would be a well-defined set of steps to convert 
> these raw data (for example, DN in an image pixel in a camera exposure 
> through a filter) into physical quantities (flux in a pixel with known 
> sky world coordinate edges and a defined spectral bandpass and start 
> and stop time).  And so we would consider the steps (bias subtraction, 
> dark subtraction, flat fielding, etc.) to take the former data and 
> produce the later as “data reduction” steps and could create a data 
> processing pipeline that performs those steps.  Anything after that 
> required knowledge of the goals of the science study and would be 
> termed “data analysis” since it was in general performed by a human. 
>  I think most historical archival data (particularly in the optical) 
> consist of either the raw data or the output of the data processing 
> pipeline with just these steps applied.
>
> However, as data have become more complex and therefore more difficult 
> for individual astronomers to perform what might have previously been 
> termed data analysis, we are seeing more of these steps being 
> performed automatically in the data processing pipelines.  The 
> increased importance of time domain astronomy, where rapid analysis of 
> data (e.g., photometry, spectra, positions etc. for detected 
> astronomical sources) is needed to trigger alerts has also contributed 
> significant momentum to this transition.  So now our data processing 
> pipelines do far more than simply converting from raw data to 
> “calibrated” data.  They also perform “data selection” steps 
> identifying the “interesting” data regions and perform various “data 
> analysis” functions such as computing draws and pdfs and selecting 
> between different models for the data based on statistical analyses. 
>  This is particularly the case for HEA where data “calibration” 
> doesn’t work the same way as in the optical and some elements of “data 
> analysis” (such as selecting appropriate source models) are required 
> to produce calibrated spectra and fluxes from raw measurements.  X-ray 
> astronomy is leading the way here largely because the instrumentation 
> now has spatial, spectral, and temporal resolution that rivals 
> ground-based optical telescopes, with effective areas and detection 
> efficiency high enough to gather significant numbers of counts for 
> X-ray bright sources.  Additionally, the detailed set of steps needed 
> to produce scientifically optimal and robust data analysis products is 
> becoming increasingly complicated and difficult for the typical 
> multi-wavelength astronomer who is not an expert in the specific 
> telescope/instrument data to perform.  These scientists are not 
> interested in becoming experts in the nuances of foibles of every 
> instrument whose data they use, but are primarily interested in the 
> final results for their science.  So there is also a strong push to 
> further automate these steps.  This is different from 2-3 decades ago, 
> where observers went to the telescope and worked directly with the 
> instrumentation, and is a result of the growth of, and access to, 
> archival data.  This lack of detailed exposure to the instrumentation 
> is not necessarily a good thing, but highlights the need for the 
> domain experts to foster the data analysis needs of the broader community.
>
> I see this merging of what were previously separate data processing, 
> data selection, and data analysis steps to accelerate even more 
> rapidly as AI enhances our capabilities in these areas.  And as I have 
> previously mentioned, I expect that in some cases the intermediate 
> data products may never be returned from the observatory.  This is the 
> future of observational astronomy, and I don’t think segregating data 
> products into these separate roles is really appropriate under these 
> circumstances.
>
> With regard to data levels or calibration levels, I do agree that each 
> project will likely have data products that fall into multiple 
> classifications.  What I don’t agree with is that our “advanced data 
> products” will always have a specific calibration level.  This is 
> similar to an image data product, where an image may be raw 
> instrumental data, in internal or a standard format, or calibrated, 
> science ready data, or enhanced such as mosaicked data, or even an 
> analysis product such as a heat map.  So if an image can be 
> calib_level 0, 1, 2, 3, or 4 we should be careful not to assume that 
> an advanced data product such as draws or pdf or region must 
> automatically have calib_level = 4.  A region data product could, for 
> example, identify the region of a detector in detector coordinates 
> that has a specific property such as region that is coated with a 
> filter material (e.g., an Al coating).  I’m not saying such a data 
> product would be queryable in ObsCore - and maybe it’s OK because 
> calib_level is an ObsCore-specific attribute - but these data product 
> types may have a wider range of applicability than just a single data 
> product level.
>
> In your example - comparing pdfs between two different data 
> collections - one would need to ensure that the pdfs represent the 
> same quantities that are similarly calibrated in the same or 
> interconvertible units (e.g., PSF-fraction corrected aperture 
> photometry of the same region of the sky in, say, optical magnitude 
> with a known flux reference and X-ray energy flux).
>
> Capturing how to handle all this appropriately and generally in the 
> next ObsCore update is going to be both important and somewhat tricky.
>
> For now, let’s keep “advanced data products” in the HEA extension 
> document.
>
> Cheers,
> —Ian
>
>> On Aug 26, 2026, at 13:20, Mireille Louys <mireille.louys at unistra.fr> 
>> wrote:
>>
>> Hi Ian, Hi HEIG readers
>>
>> Thanks for your clarifications and examples.
>>
>> Still we have written the note to sort out all the differences 
>> between the various data sets ,
>> and highlighted the benefice of the statistical validation that pdf 
>> and draws may  bring into the game.
>> I find it strange to mixe them back again , as we had in the very 
>> beginning in Nov 2025.
>>
>> The users need to segregate between the various kinds because the 3 
>> categories do not play the same role
>> in science interpretation as far as I understand.
>> This is also why data levels are defined f.i in CTAO.
>>
>> If I want to compare pdf or draws between 2 different collections , 
>> the data represented need to be fully calibrated, so
>> they have calib_level >3 at least in Obscore sense , and can be named 
>> advanced data products as in Obscore 1.1
>>
>> The term is not statisfying , but the distinction is , so we can keep 
>> this term for this note and will brainstorm later to find a better one
>> for the main Obscore extension WD ?
>>
>> This is my suggestion.
>> Best wishes , Mireille
>>
>>
>> Le 25/08/2026 à 9:22 PM, Dr. Ian N. Evans a écrit :
>>>>
>>>>
>>>> This question will need to be adressed for Obscore 1.2 , but for 
>>>> this Note we can keep "advanced data products"
>>>> with the caveat that clarification is needed.
>>>>
>>> Shall we just drop the term altogether in the document and just call 
>>> them all “data products”?
>> -- 
>> --
>> Mireille Louys, MCF (Assistant Professor)
>> Centre de données Astronomiques (CDS)       Equipe Images, ICube
>> Observatoire de Strasbourg                  Telecom Physique Strasbourg
>> 11, rue de l' Université                    300, Bd Sebastien Brandt CS 10413
>> F-67000 Strasbourg                          F-67412  Illkirch Cedex
>
>> Dr. Ian Evans
> *Astrophysicist*
> *Chandra X-ray Center*
> Center for Astrophysics | Harvard & Smithsonian
> Office: (617) 496 7846 | Cell: (617) 699 5152
> 60 Garden Street | MS 81 | Cambridge, MA 02138
>
> PastedGraphic-2.png
>
> PastedGraphic-3.png _
>
> <http://cfa.harvard.edu/>__cfa.harvard.edu 
> <http://cfa.harvard.edu/>_ | _Facebook 
> <http://cfa.harvard.edu/facebook>_ | _Twitter 
> <http://cfa.harvard.edu/twitter>_ | _YouTube 
> <http://cfa.harvard.edu/youtube>_ | _Newsletter 
> <http://cfa.harvard.edu/newsletter>_
>
>
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://mail.ivoa.net/pipermail/heig/attachments/20260901/b35c53ed/attachment-0001.htm>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: PastedGraphic-2.png
Type: image/png
Size: 581 bytes
Desc: not available
URL: <http://mail.ivoa.net/pipermail/heig/attachments/20260901/b35c53ed/attachment-0002.png>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: PastedGraphic-3.png
Type: image/png
Size: 21717 bytes
Desc: not available
URL: <http://mail.ivoa.net/pipermail/heig/attachments/20260901/b35c53ed/attachment-0003.png>


More information about the heig mailing list