MISSION 4: PAW/ROOT AND USERCODES

Aus HERMESwiki
Zur Navigation springen Zur Suche springen

Page maintainer: Andreas

Checked.png This page is considered done. It been reviewed by Caro(PAW) and David(ROOT). There may be missing elements, but they are all flagged and the text has no errors.

Your final mission is to analyze real data! To produce real physics results, you're going to need two tools. First, you need to be familiar with a good plotting program and data-manipulation package. The standard ones in particle and nuclear physics are PAW and ROOT, and you'll be learning one of those as your first task. Second, you need a usercode. This is HERMES jargon for a program that makes the bridge from ADAMO, the language in which our data is stored, to PAW or ROOT, the programs you will use to explore that data.

CERN Software Overview

PAW and ROOT are plotting and data-manipulation packages written at CERN expressly for the high-energy physics community. They are broadly similar to popular data analysis tools you may be familiar with: Mathematica, Maple, LabView, etc. However they specifically excel at tasks which high-energy needs, namely manipulation of very large data volumes, handling of complex relational data structures, and access to numerous numerical routines for fitting and statistical analysis.

PAW and ROOT are part of a huge suite of freeware code from CERN which is used by the high-energy physics community. PAW is part of CERNLIB, which is written in Fortran and has a long and distinguished history. As the name suggests, most of CERNLIB consists of code libraries. These libraries contain routines for standard numerical algorithms, memory management, particle physics simulations, and much more. xxxxxxxxxxxxx change to 3-package project structrure. PAW and ROOT are interactive programs which are basically frontends to these libraries, providing access to their functionality via a command-line interface and including graphical display capabilities. Both also have scripting languages. PAW or ROOT scripts allow you to do quite complex analysis and plotting tasks without compiling or linking anything.

Choosing PAW or ROOT

You'll have to learn PAW or ROOT ... so which one to choose? The decision depends on your coding experience:

  • Are you experienced in C++, java, or some other object-oriented language? If so, ROOT is the way to go.
  • If you have no significant experience in object-oriented programming, go with the older PAW. It's not being developed anymore but it works just fine, it's just as powerful as ROOT, and it makes no reference to object-oriented concepts.

CERNLIB, ROOT, and GEANT4 on the web

NEEDTEXT Gnome working on it The big collection of CERN xxxxxx Top-level access to CERNLIB is at this all-important link:

http://wwwasd.web.cern.ch/wwwasd/cernlib/

Learning PAW

The CERN documentation is pretty miserable but it's the only way to go, I'm afraid. To learn PAW, the best thing to do is work through the PAW tutorial at CERN, but with a goal in mind. What is coming next is a set of real-world exercises: in each, you will have a list of plots to make, and a graphics file showing what the plots should look like in the end. Keep those goals in mind as you wade through this mess of documentation!

Tour through PAW Documentation and Basics

The official PAW documentation at CERN lives here: PAW Homepage The documentation for each of the CERNLIB libraries lives here: CERNLIB documentation You should bookmark both of these pages, as the CERN IT website is notoriously difficult to navigate, with many broken links to add to the excitement.

The first page to read is the brief Fundamental PAW Objects page. It introduces the three distinct data structures which PAW deals with: histograms, vectors, and ntuples. Each type of data structure has its own virtues and its own tools and commands for manipulating and plotting the data. In becoming proficient with PAW, it's important to know which data format to use when.

  • Vectors (also called arrays) are the best data format for performing complex calculations, using commands from the powerful SIGMA package.
  • Histograms are the natural format for collecting data in bins, and are often the best format for plotting and fitting. 1D and 2D histograms are supported. The CERNLIB library HBOOK is in charge of all histogramming functions.
  • Ntuples are the best format for dealing with complex data. They contain long lists of records (e.g. one per event of interest), and each record can contain a multitude of variables (e.g. physics variables like x and Q2). One of the great powers of ntuples is that you can plot correlations between such variables very easily, and apply calculations and masks (cuts) to them. The inevitable downside of ntuples is that you have to define their structure (i.e. which variables are involved in each record) before you can use them. If the variables are all scalars, you have what's known as a row-wise ntuple. However, you can also have arrays within each record (e.g. arrays of track-level variables associated with a single event). Ntuples involving arrays are called column-wise ntuples. The HBOOK library also deals with ntuples.

An important thing to look for as you work through the PAW documentation is the commands to convert between these various formats. A very common thing to do, e.g., is to load up a histogram with data from an ntuple (thereby collecting the data into bins), then move the content of the histogram to vector format so that you can perform calculations with SIGMA.

Next, look at the PAW Components page. It's just a flowchart, but it illustrates the various libraries that are linked into PAW. Read the little text in the boxes to get an indication of what each component does, as it will help you understand the structure of the PAW commands: they are organized into menus, each of which roughly corresponds to one of these components. You can see the menu structure immediately: type paw, then after you hit return to pop open the graphics window (or enter 0 so that it doesn't pop up at all) type help at the command line prompt. You can navigate these menus to get help on every PAW command, but that's not the whole story, so ...

Onto a tour through the documentation itself.

  • The PAW Tutorial: you should work through this.
  • The PAW FAQ is amazingly useful! You will find the answers to a very large number of common PAW questions that are poorly covered in the tutorial and manual, but it's more useful once you have a bit of experience.
  • The CERNLIB documentation page has complete manuals of all the CERNLIB libraries. The main reason to visit this page at the moment is to get the full-blown PAW Users Guide which, for reasons unknown, only lives on this page, and only in gzipped postscript format. I've provided a searchable-PDF version at our HERMES web site: PAW Users Guide in PDF format. You should bookmark this page too! You can get a free hardcopy of this large manual from the DESY Computing Center in building 2. It's important to have as it contains a lot of stuff that is not in the tutorial or the online help. I'll tell you exactly which sections to pay attention to in a second.
  • The PAW Reference Manual is just an HTML version of the all the help pages within PAW.

As you read through the PAW documentation, keep your eyes open for ways to make the requested plots from the upcoming exercises!

Back to that PAW Components flowchart for a second. Here are some components to pay particular attention to, as they are rather hidden in the PAW documentation, but provide immensely useful functions.

  • SIGMA: a very powerful package for performing calculations with vectors. The best description of SIGMA is in Chapter 5 of the full PAW Users Guide
  • MINUIT: a powerful, industry-standard minimization package which is invoked by PAW's various fitting commands. It has its own manual, on the CERNLIB page, as many people link MINUIT into their usercodes and call its routines directly to perform complex fitting and minimization operations. You can skip this until you get deep into fitting.
  • COMIS: a fortran interpreter, which allows you to write little scripts in old-style Fortran 77 and call them without any compilation. With the power of Fortran at your disposal, you can perform some pretty serious calculations directly within PAW, particularly when you are looping over ntuple records. COMIS is horribly under-explained in the Users Guide, there are random references to it all over the place with no clear explanation. We'll try to solve that with one of the exercises below.
  • KUIP: this is just the user-interface which interprets the PAW commands, but it also controls the PAW scripting language which is sometimes called MACRO. You should read the following HELP pages (from within PAW) to get a feeling for this scripting syntax:
    • help functions takes you to the KUIP System Functions help page. This gives you a list of truly useful system functions that you can use to do all kinds of neat things, like string manipulation and extracting the min and max values of a histogram.
    • help syntax takes you to the 5 help pages on MACRO syntax: Expressions, Variables, Definitions, Branching, and Looping. You should read through these, they're not long.
  • HIGZ and HPLOT: these modules control the graphics aspects of PAW, which are notoriously non-intuitive. Changing fonts, choosing colors and plotting symbols, and adjusting label spacing involves cryptic 4-character graphics options and obscure numerical codes. Chapter 7 of the PAW Users Guide is the place to find the info you need. In particular, certain pages are very useful to stick on your wall as references. I've extracted them for convenience:

Your first PAW kumac

In your analysis work, you will be doing lots of plotting right in PAW: executing commands in real time to explore your data. Once you know what calculations you want to perform, however, you will want to write a script. PAW scripts are called kumacs (stands for KUip MACro) and always carry the file extension .kumac. A kumac is basically just a sequence of PAW commands, but it can also contain more complex functionality such as variables and branching (if-then-else) and looping (foreach) statements. You can write real code like this! As indicated in the previous section, check out the macro/syntax section of PAW help to learn about this scripting language.

To get you started, paste the following lines into a file called myfirst.kumac.

 a = 3
 b = 4
 c = [a] * [b]
 mess [a] times [b] is [c]
 * This is a comment and does nothing
 * We will now open a graphics file which will store all your plots
 * ... open the file on unit 22 (randomly chosen)
 for/file 22 myplot.ps
 * ... declare unit 22 as a postscript file (code -111)
 meta 22 -111
 * Create two vectors and plot them against each other
 vec/create x(5) R 1 2 3 4 5
 vec/create y(5) R 1 4 9 16 25
 * ... create an empty box with given x,y limits for our plot
 null 0 6 0 30
 * ... plot the 5 (x,y) points with a closed circle (code 20)
 symbols x y 5 20
 * Close our plot file
 for/close 22

Once you have this file saved, start PAW, hit return to pop up the graphics window, and execute your script like this:

 exec myfirst

You will now have a lovely plot file myplot.ps.

Here is another way to make exactly the same plot using the SIGMA processor, a conversion from vector to histogram format, and a histogram plotting command:

 for/file 22 myplot.ps
 meta 22 -111
 * Use SIGMA to create our two vectors
 x=$sigma(5,1#5)
 y=$sigma(x**2)
 * Create a 1D histogram (number 99) with 5 bins
 1d 99 'myhisto' 5 0 6
 * Fill this histogram with the contents of vector y
 put_vect/contents 99 y
 * Set the marker type for plotting to code 20 = filled circle
 set mtyp 20
 * Plot the histogram using this marker (NB: this command autoscales)
 h/plot 99 p

So with those two illustrations, on to the CERN PAW Tutorial, and then the exercises.

PAW exercise 1: working with vectors and SIGMA

In this example, you will not read in any data at all. Rather you will construct your own vectors, perform some calculations, and plot the result.

The plot we are going to make is an extremely important plot for doing deep-inelastic scattering physics. The plot shows the DIS kinematic regime in (x,Q2) space, as defined by the following standard cuts:

  • Q2 > 1 to make sure the virtual photon has enough spatial resolution to probe individual quarks
  • W2 > 4 to make sure we are out of the resonance region
  • y < 0.85 to minimize the size of QED radiative corrections

As there are only two independent variables in inclusive DIS at fixed beam energy, the edge of each of these constraints is a line in (x,Q2) space. Your task is to calculate and draw these lines. You will also add in these lines:

  • y < 1 reflects the fact that the virtual-photon energy cannot be greater than the beam energy; it imposes a kinematic limitation on (x,Q2) which depends on the beam energy (27.6 GeV at HERMES)
  • thetae > 40 mrad is the minimum polar scattering angle accepted by our spectrometer
  • thetae < 220 mrad is the maximum polar scattering angle accepted by our spectrometer
Your goal is to make this plot: kinema.pdf

Save it and memorize it! (You might also try different beam energies to see how the available kinematic space changes.)

As with any exercise, please give it a try first, using the techniques you learn as you go through the PAW Tutorial. Pay particular attention to the SIGMA chapter of the PAW Users Guide. Also try to figure out how to do the graphical elements, like adding labels and changing colors.

Here is one solution, in the form of a PAW kumac:

Download: kinema.kumac

Please don't look at the solution straight away: it will be much more helpful to you if you've tried it yourself for a while, give yourself a good week at least if this is you've never worked with PAW before! Once you have at least a partial solution yourself, read carefully through the solution and you'll discover a number of PAW tricks that come from experience.

PAW exercise 2: working with ntuples

In this exercise, you will read in a pre-prepared ntuple, containing DIS events that were simulated by Monte Carlo. The structure of the ntuple will be described so that you know what the variables in each record mean. Then you get your task: to make the list of requested plots. Again, if you are learning PAW for the first time, please work at the exercise for a good week before you look at the solution, or the solution probably won't make much sense! Also, there's some physics involved in this exercise, so check back with Bootcamp for help (e.g. the table of Monte Carlo particle-type codes located in the Important MC Details section.

Download: exercise.txt with all the instructions you need
Download: dis_2.rz which is the ntuple file you'll be reading in (its location on the PC farm is given in the exercise file)

The plots you are trying to make comprise several pages.

Your goal is to reproduce this series of plots: nt-AllOrdered.pdf

Once you've tried your hand at this, have a look at the solution kumac:

Download: nt.kumac

If you've given it a good try first, you'll notice some cool tricks in there. :-)

PAW exercise 3: working with COMIS functions

The fortran interpreter COMIS only speaks old-style Fortran 4, but it is a powerful way to work with really complex functions in PAW.

This example shows the Coulomb potential around four positive charges. The charges are placed in the (x,y) plane, on a square of side 5 units around the origin. The plot is a 3D picture of the Coulomb potential at each point in this plane. COMIS is so badly explained in the PAW documentation that you might as well just look at the solution straight away.

Download: fourq.kumac
Download: fourq.f

Once you have both of those files in the same directory, you can execute fourq.kumac from the PAW command line. You should get

This plot: fourq.pdf

Pretty colourful, eh. :-)


Learning ROOT

In principle ROOT is a class library similar to the CERNLIB. The core libraries include all functionalities known from the CERNLIBs with many additions that will be explained later. The front end to the ROOT libraries is a C/C++ interpreter which means that you can write macros in a programming language instead of a macro or scripting language. Moreover, C/C++ macros that run in an interactive root session can easily be compiled or code fragments simply copied to other program source codes. You therefore only have to deal with a single programming language.

Here is a list of key features and components:

  • The ROOT Dictionary allows the use of extensive runtime type informations (RTTI) for all classes, functions and globals
  • The ROOT Object I/O System allows writing objects to files and sending them via network in a machine independent way
  • BASE is a collection of base classes that are the core of ROOT
  • CONT is a collection of container classes such as arrays, vectors, list etc.
  • HIST is a collection of histogram classes including 1D, 2D and 3D histograms and graphs as well as profiles
  • MINUIT provides classes related to fitting
  • TREE is a collection of classes that allow to work with ROOT trees. A tree is the powerful ROOT version of a CERNLIB nTuple
  • MATRIX provides a linear algebra package
  • PHYSICS provides physics vector classes such as 3-vectors, 4-vectors and rotations
  • GEOM is the 3D geometry class library
  • NET is a collection of networking classes
  • GUI is a collection of GUI classes

Getting started

If you have never used ROOT before you have to set the following few environment variables:

sh:
export ROOTSYS=/opt/products/root/5.18.00
export PATH=$ROOTSYS/bin:$PATH
export LD_LIBRARY_PATH=$ROOTSYS/lib:LD_LIBRARY_PATH

or

csh:
setenv ROOTSYS /opt/products/root/5.18.00
setenv PATH ${ROOTSYS}/bin:${PATH}
setenv LD_LIBRARY_PATH ${ROOTSYS}/lib:${LD_LIBRARY_PATH}

To start an interactive ROOT session just type

$> root [-l] [-b]
  *******************************************
  *                                         *
  *        W E L C O M E  to  R O O T       *
  *                                         *
  *   Version   5.18/00   16 January 2008   *
  *                                         *
  *  You are welcome to visit our Web site  *
  *          http://root.cern.ch            *
  *                                         *
  *******************************************

ROOT 5.18/00 (trunk@21744, Jan 18 2008, 16:54:00 on linux)

CINT/ROOT C/C++ Interpreter version 5.16.29, Jan 08, 2008
Type ? for help. Commands must be C++ statements.
Enclose multiple statements between { }.

The -l option will suppress the splash screen whereas the -b option will start ROOT in batch mode. This means that all graphics output will be redirected. For running ROOT interactively on one of the batch nodes you have to use the -b option. In the interactive ROOT session you can start typing C/C++ code which will be interpreted when you hit return:

root[0] cout << "Hello World !!!" << endl;
Hello World !!!

To run a macro simply execute root with the filename of the macro

root macro.C

or from within a ROOT session via

root[0] .x macro.C

To exit an interactive ROOT session type

root[0] .q

and hit enter.

A good starting point to learn something about ROOT is the actual manual. If you are new to C++ have a look at the "A Little C++" chapter of the manual first. Here is a list of other useful chapters of the ROOT manual:

In parallel it is a good idea to work yourself through the ROOT Tutorials. From your AFS account you can find them in the directory

/opt/products/root/5.18.00/tutorials

In the following sections you will find ROOT "translations" of the PAW exercises from above. Some more analysis examples can be found on the Wiki pages for Hanna++. Some useful tricks can also be found on the ROOT Tricks page.

A ROOT example macro

Output of the example macro on the left

The following is a ROOT example macro that creates a histogram and fills it with a gaussian distribution.

1       {
2         gStyle->SetOptFit(1);
3
4         TF1 * gaussian = new TF1("gaussian", "gaus", -10, 15);
5         gaussian->SetParameters(1.0, 2.5, 3.0);
6
7         TH1F * histo = new TH1F("name", "title", 100, -10, 15);
8         histo->FillRandom("gaussian", 25000);
9         histo->Draw();
10
11        gaussian->SetParameters(1.0, 0.0, 1.0);
12        histo->Fit(gaussian);
13
14        gPad->Print("gaussian_root.png");
15      }

In line 2 the style options are changed such that fit parameters are added to the statistics box of the canvas. In lines 4 and 5 a gaussion function is created and the parameters are set (mean = 2.5, sigma = 3.0). In lines 7 - 9 a 1D histogram is created and randomly filled according to the distribution given in the function that was just created. histo->Draw() draws the histogram on the screen. In line 11 the parameters of the function are reset and in line 12 a fit to the histogram is performed. The statement in line 14 prints the picture on the screen into a png file.

ROOT version of PAW exercise 1

This exercise is an exact copy of the "PAW exercise 1: working with vectors and SIGMA". Please take the instructions from there. The ROOT result is the following:

Kinema root.png

The "official" macro can be found here: kinema.C

ROOT version of an exercise similar to PAW exercise 2

In this exercise, you will read in a prepared ROOT file, that has been created using one of the Hanna++ examples. The structure of the tree can be obtained via a TTree::Print() call:

  • EBeam - the beam energy
  • BeamCharge - the beam charge
  • Q2True - the generated Q2
  • NuTrue - the generated ν
  • xBTrue - the generated xB
  • W2True - the generated W2
  • nTracks - the number of tracks for a event
  • Tracks_ - a TClonesArray with the actual tracks

The TClonesArray is filled with objects of type BootcampTrack. Please download the header file for BootcampTrack here: BootcampTrack.hh.

For additional information on MC-related information, have a look at the Important MC Details section.

The task is to create the following plots for all reconstructed DIS events

  • The generated distributions of the DIS variables Q2, ν, xB, W2, y.
  • The reconstructed distributions of the DIS variables Q2, ν, xB, W2, y.
  • The reconstructed versus the generated distributions of the DIS variables Q2, ν, xB, W2, y.
  • The difference between generated and reconstructed versus the generated distributions of the DIS variables Q2, ν, xB, W2, y.

using the following root file.

The kinematical criteria to select DIS events can be taken from PAW exercise 1 or from the pre-defined DIS selector in the Hanna++ code environment. In addition, please ensure the scattered lepton was coming from the target cell, which means the z position of the vertex to be between 5 cm and 20 cm. (The target cell was shifted, when the Recoil Detector was installed.)

The output shall then be compared to this pdf file. If you got stuck with your attempt and need help or you achieved the task and are just curious on the "official" solution, have a look at the macro analysis1.C. This macro uses the TTree::Draw() method and a simple loop over the tree to fill the histograms. An alternative is to use a selector class derived from TSelector. This is the ROOT recommended solution for analyzing trees. However, for such a simple example probably not suitable. The solution using TSelector can be downloaded here: analysis2.h and analysis2.C

ROOT version of PAW exercise 3

SixQ root.png

This exercise is an extended ROOT version of PAW exercise 3. In the example the Coulomb potential of six positive charges arranged around the origin is plotted as a 2D function. The macro will create a canvas with 4 pads. In each pad the function is drawn in different styles. The result of the exercise is shown in the figure on the right. The macro used to produce this figure can be downloaded here: sixq.C.

In ROOT, functions (1D, 2D and 3D) are objects of type TF1, TF2 or TF3. The actual function can either be an interpreted formula, an interpreted function in a macro or a compiled function. A trivial example for an interpreted formula is [0]+[1]*x+[2]*x*x where [n] are function parameters. For details on how to use the ROOT function classes have a look at the class documentations for TF1, TF2 and TF3. In addition ROOT comes with some nice tutorials concerning functions: formula1.C and myfit.C.

Overview of Usercodes

A usercode is a program you write yourself that reads in our ADAMO data files, performs calculations and filtering, and converts the info you want into a format that PAW or ROOT can read.