[Heig] HEIG Obscore Note - ready for final review/ comments on the Note #part 1

Dr. Ian N. Evans ievans at cfa.harvard.edu
Thu Aug 27 17:33:51 CEST 2026


Hi Mireille,

I think this is an area where we might disagree ;-)

In my view, these “advanced” data product types can play roles in all of the 3 categories (by which I assume you are meaning data processing, data selection, and data analysis) and so segregation by role is difficult.

Historically we used to acquire raw images, spectra, etc. at the telescope and there would be a well-defined set of steps to convert these raw data (for example, DN in an image pixel in a camera exposure through a filter) into physical quantities (flux in a pixel with known sky world coordinate edges and a defined spectral bandpass and start and stop time).  And so we would consider the steps (bias subtraction, dark subtraction, flat fielding, etc.) to take the former data and produce the later as “data reduction” steps and could create a data processing pipeline that performs those steps.  Anything after that required knowledge of the goals of the science study and would be termed “data analysis” since it was in general performed by a human.  I think most historical archival data (particularly in the optical) consist of either the raw data or the output of the data processing pipeline with just these steps applied.

However, as data have become more complex and therefore more difficult for individual astronomers to perform what might have previously been termed data analysis, we are seeing more of these steps being performed automatically in the data processing pipelines.  The increased importance of time domain astronomy, where rapid analysis of data (e.g., photometry, spectra, positions etc. for detected astronomical sources) is needed to trigger alerts has also contributed significant momentum to this transition.  So now our data processing pipelines do far more than simply converting from raw data to “calibrated” data.  They also perform “data selection” steps identifying the “interesting” data regions and perform various “data analysis” functions such as computing draws and pdfs and selecting between different models for the data based on statistical analyses.  This is particularly the case for HEA where data “calibration” doesn’t work the same way as in the optical and some elements of “data analysis” (such as selecting appropriate source models) are required to produce calibrated spectra and fluxes from raw measurements.  X-ray astronomy is leading the way here largely because the instrumentation now has spatial, spectral, and temporal resolution that rivals ground-based optical telescopes, with effective areas and detection efficiency high enough to gather significant numbers of counts for X-ray bright sources.  Additionally, the detailed set of steps needed to produce scientifically optimal and robust data analysis products is becoming increasingly complicated and difficult for the typical multi-wavelength astronomer who is not an expert in the specific telescope/instrument data to perform.  These scientists are not interested in becoming experts in the nuances of foibles of every instrument whose data they use, but are primarily interested in the final results for their science.  So there is also a strong push to further automate these steps.  This is different from 2-3 decades ago, where observers went to the telescope and worked directly with the instrumentation, and is a result of the growth of, and access to, archival data.  This lack of detailed exposure to the instrumentation is not necessarily a good thing, but highlights the need for the domain experts to foster the data analysis needs of the broader community.

I see this merging of what were previously separate data processing, data selection, and data analysis steps to accelerate even more rapidly as AI enhances our capabilities in these areas.  And as I have previously mentioned, I expect that in some cases the intermediate data products may never be returned from the observatory.  This is the future of observational astronomy, and I don’t think segregating data products into these separate roles is really appropriate under these circumstances.

With regard to data levels or calibration levels, I do agree that each project will likely have data products that fall into multiple classifications.  What I don’t agree with is that our “advanced data products” will always have a specific calibration level.  This is similar to an image data product, where an image may be raw instrumental data, in internal or a standard format, or calibrated, science ready data, or enhanced such as mosaicked data, or even an analysis product such as a heat map.  So if an image can be calib_level 0, 1, 2, 3, or 4 we should be careful not to assume that an advanced data product such as draws or pdf or region must automatically have calib_level = 4.  A region data product could, for example, identify the region of a detector in detector coordinates that has a specific property such as region that is coated with a filter material (e.g., an Al coating).  I’m not saying such a data product would be queryable in ObsCore - and maybe it’s OK because calib_level is an ObsCore-specific attribute - but these data product types may have a wider range of applicability than just a single data product level.

In your example - comparing pdfs between two different data collections - one would need to ensure that the pdfs represent the same quantities that are similarly calibrated in the same or interconvertible units (e.g., PSF-fraction corrected aperture photometry of the same region of the sky in, say, optical magnitude with a known flux reference and X-ray energy flux).

Capturing how to handle all this appropriately and generally in the next ObsCore update is going to be both important and somewhat tricky.

For now, let’s keep “advanced data products” in the HEA extension document.

Cheers,
—Ian

> On Aug 26, 2026, at 13:20, Mireille Louys <mireille.louys at unistra.fr> wrote:
> 
> Hi Ian, Hi HEIG readers 
> 
> Thanks for your clarifications and examples.
> 
> Still we have written the note to sort out all the differences between the various data sets , 
> and highlighted the benefice of the statistical validation that pdf and draws may  bring into the game. 
> I find it strange to mixe them back again , as we had in the very beginning in Nov 2025. 
> 
> The users need to segregate between the various kinds because the 3 categories do not play the same role
> in science interpretation as far as I understand. 
> This is also why data levels are defined f.i in CTAO. 
> 
> If I want to compare pdf or draws between 2 different collections , the data represented need to be fully calibrated, so 
> they have calib_level >3 at least in Obscore sense , and can be named advanced data products as in Obscore 1.1 
> 
> The term is not statisfying , but the distinction is , so we can keep this term for this note and will brainstorm later to find a better one 
> for the main Obscore extension WD ? 
> 
> This is my suggestion. 
> Best wishes , Mireille
> 
> 
> 
> 
> Le 25/08/2026 à 9:22 PM, Dr. Ian N. Evans a écrit :
>>> 
>>> This question will need to be adressed for Obscore 1.2 , but for this Note we can keep "advanced data products" 
>>> with the caveat that clarification is needed. 
>>> 
>> Shall we just drop the term altogether in the document and just call them all “data products”?
> -- 
> --
> Mireille Louys, MCF (Assistant Professor)
> Centre de données Astronomiques (CDS)       Equipe Images, ICube
> Observatoire de Strasbourg                  Telecom Physique Strasbourg
> 11, rue de l' Université                    300, Bd Sebastien Brandt CS 10413
> F-67000 Strasbourg                          F-67412  Illkirch Cedex

—

Dr. Ian Evans
Astrophysicist
Chandra X-ray Center
Center for Astrophysics | Harvard & Smithsonian

Office: (617) 496 7846 | Cell: (617) 699 5152
60 Garden Street | MS 81 | Cambridge, MA 02138



 


 <http://cfa.harvard.edu/>cfa.harvard.edu <http://cfa.harvard.edu/> | Facebook <http://cfa.harvard.edu/facebook> | Twitter <http://cfa.harvard.edu/twitter> | YouTube <http://cfa.harvard.edu/youtube> | Newsletter <http://cfa.harvard.edu/newsletter>
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://mail.ivoa.net/pipermail/heig/attachments/20260827/1fca7c32/attachment-0001.htm>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: PastedGraphic-2.png
Type: image/png
Size: 581 bytes
Desc: not available
URL: <http://mail.ivoa.net/pipermail/heig/attachments/20260827/1fca7c32/attachment-0002.png>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: PastedGraphic-3.png
Type: image/png
Size: 21717 bytes
Desc: not available
URL: <http://mail.ivoa.net/pipermail/heig/attachments/20260827/1fca7c32/attachment-0003.png>


More information about the heig mailing list