ho22joshua commited on
Commit
a900118
·
1 Parent(s): 1eecff3

docs: normalize inline notation in model card

Browse files
Files changed (1) hide show
  1. README.md +2 -2
README.md CHANGED
@@ -85,7 +85,7 @@ This methodological framework demonstrates the potential of foundation models to
85
 
86
  ## GNN Architecture
87
 
88
- We implement a Graph Neural Network (GNN) architecture that naturally accommodates the point-cloud structure of particle physics data, employing the `DGL` framework with a `PyTorch` backend (Wang et al. 2019; Paszke et al. 2019). The GNN naturally handles graphs of varying node and edge counts through the message-passing framework, without requiring padding or truncation, making it well suited to collision events where the number of reconstructed objects varies from event to event. A fully connected graph is constructed for each event, with nodes corresponding to reconstructed jets, electrons, muons, photons, and \\(\vec{E}_T^{\text{miss}}\\). The features of each node include the four-momentum \\((p_T,\eta,\phi,E)\\) of the object with a massless assumption (\\(E=p_T\cosh\eta\\)), the b-tagging label (for jets), the charge (for leptons), and an integer labeling the type of object represented by the node. We use a placeholder value of 0 for features which are not defined for every node type such as the b-jet tag, lepton charge, or the pseudorapidity of \\(\vec{E}_T^{\text{miss}}\\). An explicit masking mechanism was not tested in this work and represents a potential avenue for future improvement. We assign the angular distances (\\(\Delta \eta, \Delta \phi, \Delta R\\)) as edge features and the number of nodes \\(N\\) in the graph as a global feature. We denote the node features \\(\{\vec x_i\}\\), edge features \\(\{\vec y_{ij}\}\\), and global features \\(\{\vec z\}\\).
89
 
90
  The GNN model is based on the graph network architecture described in (Battaglia et al. 2018) using simple multilayer perceptron (MLP) feature functions and summation aggregation. The model is comprised of three primary components: an encoder, the graph network, and a decoder. In the encoder, three MLPs embed the nodes, edges, and global features into a latent space of dimension 64. The graph network block, which is designed to facilitate message passing between different domains of the graph, performs an edge update \\(f_e\\), followed by a node update \\(f_n\\), and finally a global update \\(f_g\\), all defined below. The inputs to each update MLP are concatenated.
91
 
@@ -135,7 +135,7 @@ This approach combines both classification and regression tasks to characterize
135
 
136
  We develop a comprehensive set of 41 labels that capture both particle multiplicities and kinematic properties. This approach increases prediction granularity and enhances model interpretability. By training the model to predict event kinematics rather than event identification, we create a task-independent framework that can potentially generalize better to novel scenarios not seen during pretraining.
137
 
138
- The particle multiplicity labels count the number of Higgs bosons (\\(n_{\text{higgs}}\\)), top quarks (\\(n_{\text{tops}}\\)), vector bosons (\\(n_V\\)), W bosons (\\(n_W\\)), and Z bosons (\\(n_Z\\)). The kinematic labels characterize the transverse momentum (\\(p_T\\)), pseudorapidity (\\(\eta\\)), and azimuthal angle (\\(\phi\\)) of Higgs bosons and top quarks through binned classifications.
139
 
140
  For Higgs bosons, \\(p_T\\) is categorized into three ranges: (0, 30) GeV, (30, 200) GeV, and (200, \\(\infty\\)) GeV, with the upper range particularly sensitive to potential BSM effects. Similarly, both leading and subleading top quarks have \\(p_T\\) classifications spanning (0, 30) GeV, (30, 300) GeV, and (300, \\(\infty\\)) GeV. When no particle exists within a specific \\(p_T\\) range, the corresponding label is set to \\([0, 0, 0]\\). For all particles, \\(\eta\\) measurements are divided into 4 bins with boundaries at \\([-1.5, 0, 1.5]\\), while \\(\phi\\) measurements use 4 bins with boundaries at \\([-\frac{\pi}{2}, 0, \frac{\pi}{2}]\\). As with \\(p_T\\), both \\(\eta\\) and \\(\phi\\) labels default to \\([0, 0, 0, 0]\\) in the absence of a particle. This comprehensive labeling schema enables fine-grained learning of kinematic distributions and particle multiplicities, essential for characterizing complex collision events.
141
 
 
85
 
86
  ## GNN Architecture
87
 
88
+ We implement a Graph Neural Network (GNN) architecture that naturally accommodates the point-cloud structure of particle physics data, employing the DGL framework with a PyTorch backend (Wang et al. 2019; Paszke et al. 2019). The GNN naturally handles graphs of varying node and edge counts through the message-passing framework, without requiring padding or truncation, making it well suited to collision events where the number of reconstructed objects varies from event to event. A fully connected graph is constructed for each event, with nodes corresponding to reconstructed jets, electrons, muons, photons, and the missing transverse-energy vector `E⃗_T^miss`. The features of each node include the four-momentum `(pT, η, ϕ, E)` of the object with a massless assumption (`E = pT cosh η`), the b-tagging label (for jets), the charge (for leptons), and an integer labeling the type of object represented by the node. We use a placeholder value of 0 for features which are not defined for every node type such as the b-jet tag, lepton charge, or the pseudorapidity of `E⃗_T^miss`. An explicit masking mechanism was not tested in this work and represents a potential avenue for future improvement. We assign the angular distances `(Δη, Δϕ, ΔR)` as edge features and the number of nodes `N` in the graph as a global feature. We denote the node features `{x⃗_i}`, edge features `{y⃗_ij}`, and global features `{z⃗}`.
89
 
90
  The GNN model is based on the graph network architecture described in (Battaglia et al. 2018) using simple multilayer perceptron (MLP) feature functions and summation aggregation. The model is comprised of three primary components: an encoder, the graph network, and a decoder. In the encoder, three MLPs embed the nodes, edges, and global features into a latent space of dimension 64. The graph network block, which is designed to facilitate message passing between different domains of the graph, performs an edge update \\(f_e\\), followed by a node update \\(f_n\\), and finally a global update \\(f_g\\), all defined below. The inputs to each update MLP are concatenated.
91
 
 
135
 
136
  We develop a comprehensive set of 41 labels that capture both particle multiplicities and kinematic properties. This approach increases prediction granularity and enhances model interpretability. By training the model to predict event kinematics rather than event identification, we create a task-independent framework that can potentially generalize better to novel scenarios not seen during pretraining.
137
 
138
+ The particle multiplicity labels count the number of Higgs bosons (`n_higgs`), top quarks (`n_tops`), vector bosons (`n_V`), W bosons (`n_W`), and Z bosons (`n_Z`). The kinematic labels characterize the transverse momentum (`pT`), pseudorapidity (`η`), and azimuthal angle (`ϕ`) of Higgs bosons and top quarks through binned classifications.
139
 
140
  For Higgs bosons, \\(p_T\\) is categorized into three ranges: (0, 30) GeV, (30, 200) GeV, and (200, \\(\infty\\)) GeV, with the upper range particularly sensitive to potential BSM effects. Similarly, both leading and subleading top quarks have \\(p_T\\) classifications spanning (0, 30) GeV, (30, 300) GeV, and (300, \\(\infty\\)) GeV. When no particle exists within a specific \\(p_T\\) range, the corresponding label is set to \\([0, 0, 0]\\). For all particles, \\(\eta\\) measurements are divided into 4 bins with boundaries at \\([-1.5, 0, 1.5]\\), while \\(\phi\\) measurements use 4 bins with boundaries at \\([-\frac{\pi}{2}, 0, \frac{\pi}{2}]\\). As with \\(p_T\\), both \\(\eta\\) and \\(\phi\\) labels default to \\([0, 0, 0, 0]\\) in the absence of a particle. This comprehensive labeling schema enables fine-grained learning of kinematic distributions and particle multiplicities, essential for characterizing complex collision events.
141