Skip to content
Merged
Show file tree
Hide file tree
Changes from 4 commits
Commits
Show all changes
67 commits
Select commit Hold shift + click to select a range
3469368
Fixing loaded_commondata_with_cuts import and add theory_covmat to ex…
andreab1997 Feb 22, 2022
80bba9d
Fixed loading of theory_covmat also for user provided covmat
andreab1997 Feb 24, 2022
815a595
Fixed theory_covmat flags
andreab1997 Feb 24, 2022
539f366
Fixed some doc
andreab1997 Mar 1, 2022
99ffdd4
removed old runcards and added new working one
andreab1997 Mar 10, 2022
f1ac85b
Fixed conflicts
andreab1997 Mar 10, 2022
6e64443
Fixing conflicts
andreab1997 Mar 10, 2022
2601ce6
Removing comments
andreab1997 Mar 10, 2022
9a8cf90
Fixing stuffs
andreab1997 Mar 10, 2022
88adcad
Changing n3fit_exec
andreab1997 Mar 10, 2022
3c0d0f5
Starting to include thcovmat in make_replica
andreab1997 Mar 15, 2022
6278c58
Fixing
andreab1997 Mar 15, 2022
6a2c0e7
Added sqrt of thcovmat to make_replica
andreab1997 Mar 15, 2022
2b61504
Fixed sqrt of covmat
andreab1997 Mar 15, 2022
73e578c
Fixing number for sqrt
andreab1997 Mar 15, 2022
3789d1e
Fixing loading of thcovmat
andreab1997 Mar 16, 2022
82ce369
First solution to sqrt of thcovmat
andreab1997 Mar 16, 2022
234cf80
Added theory_covmat to additive contrib for make_replica
andreab1997 Mar 17, 2022
9d83d75
Added flags for thcovmat
andreab1997 Mar 17, 2022
2064796
Adding regularization to thcovmat
andreab1997 Mar 17, 2022
abc1a0d
Fixing flags
andreab1997 Mar 17, 2022
9dc505c
Fixed make_replica and vp-comparefits
andreab1997 Mar 18, 2022
dd96091
Remove pdb
andreab1997 Mar 18, 2022
36b0a55
Changing implementation of make_replica (1st step)
andreab1997 Mar 18, 2022
c3dc6d7
Implemented make_replica with full covmat
andreab1997 Mar 18, 2022
9fe8d25
Added regularization to thcovmat
andreab1997 Mar 18, 2022
91a9c80
Removing a pdb
andreab1997 Mar 19, 2022
43e60ee
Fix wrong t0 for make_replica
andreab1997 Mar 21, 2022
662bcf0
Added flags for t0
andreab1997 Mar 21, 2022
0b070c4
minor changes
andreab1997 Mar 24, 2022
f1c514a
Separated loops in make_replica
andreab1997 Mar 24, 2022
d43b357
solved bug in thcovmat order
andreab1997 Mar 24, 2022
7e89de0
Fixed bug in pseudodata
andreab1997 Mar 25, 2022
bf4f841
Minor change
andreab1997 Mar 25, 2022
684cba5
Restoring possibility of separating mult errors for replica generation
andreab1997 Mar 25, 2022
0497427
Providing defaults
andreab1997 Mar 25, 2022
f52c492
Changed flags names and added comments
andreab1997 Mar 25, 2022
d615806
Fixing flags and covmats
andreab1997 Mar 25, 2022
585e69a
Changed default
andreab1997 Mar 25, 2022
9a1e9f6
Added new sampling flag in runcards
andreab1997 Mar 28, 2022
166873c
Starting fix of tests
andreab1997 Mar 30, 2022
05d3d61
Fixed test_pythonmakereplica
andreab1997 Mar 31, 2022
419e9b0
Fixed test_pseudodata
andreab1997 Mar 31, 2022
0ce016b
Fixed regressions test
andreab1997 Mar 31, 2022
894138a
Added docs
andreab1997 Mar 31, 2022
c94f741
Fixed test_fit in n3fit
andreab1997 Mar 31, 2022
e9d81f8
Fixed test_fit_and_timing
andreab1997 Mar 31, 2022
176eb3d
Resolve conflicting files
andreab1997 Mar 31, 2022
a4f4133
minor changes
andreab1997 Mar 31, 2022
89ac78f
Added produce action in config
andreab1997 Apr 5, 2022
793de60
Fixed tests
andreab1997 Apr 5, 2022
fce0ad8
Added datasets to pythonmakereplica tests
andreab1997 Apr 5, 2022
9007322
Minor changes and docs
andreab1997 Apr 6, 2022
99185dd
Minor changes
andreab1997 Apr 6, 2022
1d7b089
Restored tests and commondata files to master version
andreab1997 Apr 6, 2022
b7cf514
Merge branch 'master' into fix_thcovmat_fit
andreab1997 Apr 7, 2022
1d68e85
Fixing conflict
andreab1997 May 30, 2022
a922e26
fixed conflict again
andreab1997 May 30, 2022
62392de
Merge branch 'master' into fix_thcovmat_fit
andreab1997 May 30, 2022
7d8f4e8
Rerunning the tests
andreab1997 May 30, 2022
bbd3bc0
Merge branch 'fix_thcovmat_fit' of github.com:NNPDF/nnpdf into fix_th…
andreab1997 May 30, 2022
d84bb2f
Some minor corrections
andreab1997 Jul 1, 2022
e3d41be
Other minor changes
andreab1997 Jul 1, 2022
d4b4702
Final fixes
andreab1997 Jul 4, 2022
4d768f0
Reformatting funcs
andreab1997 Jul 4, 2022
1c5205a
removed file control
andreab1997 Jul 6, 2022
a10450e
Implemented check of files
andreab1997 Jul 6, 2022
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion doc/sphinx/source/vp/theorycov/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Summary
- Theoretical covariance matrices are built according to the various prescriptions
in :ref:`prescrips`.

- The prescription must be one of 3 point, 5 point, 5bar point, 7 point or 9 point. You can specify
- The prescription must be one of 3 point, 3r point, 3f point, 5 point, 5bar point, 7 point or 9 point. You can specify
this using ``point_prescription: "x point"`` in the runcard. The translation of this flag
into the relevant ``theoryids`` is handled by the ``scalevariations`` module in ``validphys``.

Expand Down
5 changes: 0 additions & 5 deletions doc/sphinx/source/vp/theorycov/runcard_layout.rst
Original file line number Diff line number Diff line change
@@ -1,11 +1,6 @@
Important information about runcard layout
==========================================

- The flag ``fivetheories`` specifies the choice of 5 or
:math:`\bar{5}` prescription for the case of 5 input theories. You
must assign a value ``nobar`` or ``bar`` correspondingly. If you do
not do this, ``validphys`` will give an error.

- The default behaviour for the 7-point prescription is to use Gavin
Salam's modification to it. To use the original 7-point prescription
instead, the ``seventheories`` flag must be set to ``original``.
Expand Down
3 changes: 1 addition & 2 deletions n3fit/src/n3fit/performfit.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ def performfit(
tensorboard=None,
debug=False,
maxcores=None,
parallel_models=False
parallel_models=False,
):
"""
This action will (upon having read a validcard) process a full PDF fit
Expand Down Expand Up @@ -148,7 +148,6 @@ def performfit(
# [list of all NN seeds]
# )
#

n_models = len(replicas_nnseed_fitting_data_dict)
if parallel_models and n_models != 1:
replicas, replica_experiments, nnseeds = zip(*replicas_nnseed_fitting_data_dict)
Expand Down
8 changes: 7 additions & 1 deletion n3fit/src/n3fit/scripts/n3fit_exec.py
Original file line number Diff line number Diff line change
Expand Up @@ -145,7 +145,13 @@ def from_yaml(cls, o, *args, **kwargs):
validation_action = namespace + "validation_pseudodata"

N3FIT_FIXED_CONFIG['actions_'].extend((training_action, validation_action))

N3FIT_FIXED_CONFIG['theory_covmat_flag'] = False
N3FIT_FIXED_CONFIG['use_user_uncertainties'] = None
N3FIT_FIXED_CONFIG['use_scalevar_uncertainties'] = None
Comment thread
scarlehoff marked this conversation as resolved.
Outdated
if file_content.get('theorycovmatconfig') is not None:
N3FIT_FIXED_CONFIG['theory_covmat_flag'] = True
N3FIT_FIXED_CONFIG['use_user_uncertainties'] = file_content.get('theorycovmatconfig').get('use_user_uncertainties')
N3FIT_FIXED_CONFIG['use_scalevar_uncertainties'] = file_content.get('theorycovmatconfig').get('use_scalevar_uncertainties')
Comment thread
andreab1997 marked this conversation as resolved.
Outdated
file_content.update(N3FIT_FIXED_CONFIG)
return cls(file_content, *args, **kwargs)

Expand Down
1 change: 0 additions & 1 deletion n3fit/src/n3fit/scripts/vp_setupfit.py
Original file line number Diff line number Diff line change
Expand Up @@ -158,7 +158,6 @@ def from_yaml(cls, o, *args, **kwargs):
SETUPFIT_FIXED_CONFIG['actions_'] += [filter_action]
else:
SETUPFIT_FIXED_CONFIG['actions_'] += [check_n3fit_action, filter_action]

if file_content.get('theorycovmatconfig') is not None:
SETUPFIT_FIXED_CONFIG['actions_'].append(
'datacuts::theory::theorycovmatconfig nnfit_theory_covmat')
Expand Down
8 changes: 6 additions & 2 deletions validphys2/src/validphys/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -1053,11 +1053,9 @@ def produce_nnfit_theory_covmat(
# Only user uncertainties
from validphys.theorycovariance.construction import user_covmat_fitting
f = user_covmat_fitting

Comment thread
andreab1997 marked this conversation as resolved.
@functools.wraps(f)
Comment thread
andreab1997 marked this conversation as resolved.
def res(*args, **kwargs):
return f(*args, **kwargs)

Comment thread
andreab1997 marked this conversation as resolved.
# Set this to get the same filename regardless of the action.
res.__name__ = "theory_covmat"
return res
Expand Down Expand Up @@ -1504,6 +1502,11 @@ def produce_group_dataset_inputs_by_metadata(
{"data_input": NSList(group, nskey="dataset_input"), "group_name": name}
for name, group in res.items()
]

def produce_group_dataset_inputs_by_fitting_group(self, data_input, theory_covmat_flag):
Comment thread
andreab1997 marked this conversation as resolved.
Outdated
if theory_covmat_flag is True:
Comment thread
andreab1997 marked this conversation as resolved.
Outdated
return self.produce_group_dataset_inputs_by_metadata(data_input, "custom_group")
Comment thread
andreab1997 marked this conversation as resolved.
Outdated
return self.produce_group_dataset_inputs_by_metadata(data_input, "experiment")


def produce_group_dataset_inputs_by_experiment(self, data_input):
Expand All @@ -1512,6 +1515,7 @@ def produce_group_dataset_inputs_by_experiment(self, data_input):
def produce_group_dataset_inputs_by_process(self, data_input):
return self.produce_group_dataset_inputs_by_metadata(data_input, "nnpdf31_process")


def produce_scale_variation_theories(self, theoryid, point_prescription):
"""Produces a list of theoryids given a theoryid at central scales and a point
prescription. The options for the latter are '3 point', '5 point', '5bar point', '7 point'
Expand Down
49 changes: 48 additions & 1 deletion validphys2/src/validphys/covmats.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@
import numpy as np
import pandas as pd
import scipy.linalg as la
import pathlib

from reportengine import collect
from reportengine.table import table
Expand All @@ -23,7 +24,7 @@
from validphys.core import PDF, DataGroupSpec, DataSetSpec
from validphys.covmats_utils import construct_covmat, systematics_matrix
from validphys.results import ThPredictionsResult

from validphys.commondata import loaded_commondata_with_cuts
Comment thread
andreab1997 marked this conversation as resolved.
log = logging.getLogger(__name__)
Comment thread
andreab1997 marked this conversation as resolved.

INTRA_DATASET_SYS_NAME = ("UNCORR", "CORR", "THEORYUNCORR", "THEORYCORR")
Expand Down Expand Up @@ -227,6 +228,17 @@ def dataset_inputs_covmat_from_systematics(
covmat,
norm_threshold=norm_threshold
)
# try:
# theory_covmat_path = pathlib.Path.cwd()
# data = pd.read_csv(theory_covmat_path / "prov_moredata" / "tables" / "datacuts_theory_theorycovmatconfig_theory_covmat_custom.csv", sep='\t')
# datael = data.iloc[3:]
# datael = datael.drop(['group'], axis=1)
# datael = datael.drop(['Unnamed: 1'], axis=1)
# datael = datael.drop(['Unnamed: 2'], axis=1)
# theory_covmat = np.copy(datael.values)
#except FileNotFoundError:
# theory_covmat = np.zeros(covmat.shape)
#total_covmat = np.add(covmat, theory_covmat)
Comment thread
andreab1997 marked this conversation as resolved.
Outdated
return covmat


Expand Down Expand Up @@ -339,6 +351,41 @@ def dataset_inputs_t0_covmat_from_systematics(
_list_of_central_values=dataset_inputs_t0_predictions
)

def dataset_inputs_t0_total_covmat(dataset_inputs_loaded_cd_with_cuts,
*,
data_input,
use_weights_in_covmat=True,
norm_threshold=None,
dataset_inputs_t0_predictions,
output_path,
theory_covmat_flag,
use_user_uncertainties,
use_scalevar_uncertainties):
exp_covmat = dataset_inputs_covmat_from_systematics(
dataset_inputs_loaded_cd_with_cuts,
data_input,
use_weights_in_covmat,
norm_threshold=norm_threshold,
_list_of_central_values=dataset_inputs_t0_predictions
)
if theory_covmat_flag is True:
generic_path = None
if use_scalevar_uncertainties is True:
if use_user_uncertainties is True:
generic_path = pathlib.Path("datacuts_theory_theorycovmatconfig_total_theory_covmat.csv")
else:
generic_path = pathlib.Path("datacuts_theory_theorycovmatconfig_theory_covmat_custom.csv")
else:
if use_user_uncertainties is True:
generic_path = pathlib.Path("datacuts_theory_theorycovmatconfig_user_covmat.csv")
else:
generic_path = pathlib.Path("datacuts_theory_theorycovmatconfig_theory_covmat_custom.csv")
theorypath = pathlib.Path(str(output_path/"tables"/generic_path.relative_to(generic_path.anchor)))
Comment thread
andreab1997 marked this conversation as resolved.
Outdated
theory_covmat = pd.read_csv(theorypath, sep='\t')
theory_covmat = theory_covmat.iloc[3:].drop(['group'], axis=1).drop(['Unnamed: 1'], axis=1).drop(['Unnamed: 2'], axis=1)
return np.add(exp_covmat,theory_covmat.values.astype(np.float))
return exp_covmat

@scarlehoff scarlehoff Mar 23, 2022

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think for these two functions (and all the branches they create) we need to rethink a bit the names and what we want to have.

Imho we should have sampling_total_covmat and fitting_total_covmat (or something like that) so that it's clear what they should be used for.

Personally I would just enforce that sampling does not include t0 and fitting does but yesterday talking with @Radonirinaunimi he was interested in having the possibility of setting that flag on and off so we might as well included (since you already did).

We can sketch the different names / branches that we want in a blackboard when I'm at the office, it will be faster.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok sure. Till then I leave the names as they are now

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

BTW I believe that we should leave the flags for t0 as well

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes! It would help if such a flag is available, at least from a conceptual point of view.



def sqrt_covmat(covariance_matrix):
"""Function that computes the square root of the covariance matrix.
Expand Down
7 changes: 3 additions & 4 deletions validphys2/src/validphys/n3fit_data.py
Original file line number Diff line number Diff line change
Expand Up @@ -188,7 +188,7 @@ def _mask_fk_tables(dataset_dicts, tr_masks):
def fitting_data_dict(
data,
make_replica,
dataset_inputs_t0_covmat_from_systematics,
dataset_inputs_t0_total_covmat,
tr_masks,
kfold_masks,
diagonal_basis=None,
Expand Down Expand Up @@ -244,7 +244,7 @@ def fitting_data_dict(
datasets = common_data_reader_experiment(spec_c, data)

# t0 covmat
covmat = dataset_inputs_t0_covmat_from_systematics
covmat = dataset_inputs_t0_total_covmat
inv_true = np.linalg.inv(covmat)

if diagonal_basis:
Expand Down Expand Up @@ -297,7 +297,6 @@ def fitting_data_dict(
folds["training"].append(fold[tr_mask])
folds["validation"].append(fold[vl_mask])
folds["experimental"].append(~fold)

dict_out = {
"datasets": datasets_copy,
"name": str(data),
Expand All @@ -320,7 +319,7 @@ def fitting_data_dict(
}
return dict_out

exps_fitting_data_dict = collect("fitting_data_dict", ("group_dataset_inputs_by_experiment",))
exps_fitting_data_dict = collect("fitting_data_dict", ("group_dataset_inputs_by_fitting_group",))

def replica_nnseed_fitting_data_dict(replica, exps_fitting_data_dict, replica_nnseed):
"""For a single replica return a tuple of the inputs to this function.
Expand Down