Skip to main content

Submit Functional Queries to a ServiceX endpoint.

Project description

func_adl_servicex

Send func_adl expressions to a ServiceX endpoint

GitHub Actions Status Code Coverage

PyPI version Supported Python versions

Introduction

This package contains the single object ServiceXSourceXAOD and ``ServiceXSourceUpROOTwhich can be used as a root of afunc_adl` expression to query large LHC datasets from an active `ServiceX` instance located on the net.

See below for simple examples.

Further Information

  • servicex documentation
  • func_adl documentation

Usage

To use func_adl on servicex, the only func_adl package you only need to install this package. All others required will be pulled in as dependencies of this package.

Using the xAOD backend

See the further information for documentation above to understand how this works. Here is a quick sample that will run against an ATLAS xAOD backend in servicex to get out jet pt's for those jets with pt > 30 GeV.

from func_adl_servicex import ServiceXSourceXAOD

dataset_xaod = "mc15_13TeV:mc15_13TeV.361106.PowhegPythia8EvtGen_AZNLOCTEQ6L1_Zee.merge.DAOD_STDM3.e3601_s2576_s2132_r6630_r6264_p2363_tid05630052_00"
ds = ServiceXSourceXAOD(dataset_xaod)
data = ds \
    .SelectMany('lambda e: (e.Jets("AntiKt4EMTopoJets"))') \
    .Where('lambda j: (j.pt()/1000)>30') \
    .Select('lambda j: j.pt()') \
    .AsAwkwardArray(["JetPt"]) \
    .value()

print(data['JetPt'])

Using the CMS Run 1 AOD backend

See the further information for documentation above to understand how this works. Here is a quick sample that will run against an CMS Run 1 AOD backend in servicex. It turns against a 6 TB CMS Open Data dataset, selecting global muons with a pT greater than 30 GeV.

from func_adl_servicex import ServiceXSourceCMSRun1AOD

dataset_xaod = "cernopendata://16"
ds = ServiceXSourceCMSRun1AOD(dataset_xaod)
data = ds \
data = ServiceXSourceCMSRun1AOD("cernopendata://16") \
    .SelectMany(lambda e: e.TrackMuons("globalMuons")) \
    .Where(lambda m: m.pt() > 30) \
    .Select(lambda m: m.pt()) \
    .AsAwkwardArray(['mu_pt']) \
    .value()

print(data['mu_pt'])

Using the uproot backend

See the further information for documentation above to understand how this works. Here is a quick sample that will run against a ROOT file (TTree) in the uproot backend in servicex to get out jet pt's. Note that the image name tag is likely wrong here. See XXX to get the current one.

from servicex import ServiceXDataset
from func_adl_servicex import ServiceXSourceUpROOT


dataset_uproot = "user.kchoi:user.kchoi.ttHML_80fb_ttbar"
uproot_transformer_image = "sslhep/servicex_func_adl_uproot_transformer:issue6"

sx_dataset = ServiceXDataset(dataset_uproot, image=uproot_transformer_image)
ds = ServiceXSourceUpROOT(sx_dataset, "nominal")
data = ds.Select("lambda e: {'lep_pt_1': e.lep_Pt_1, 'lep_pt_2': e.lep_Pt_2}") \
    .AsParquetFiles('junk.parquet') \
    .value()

print(data)

Running on Local Datasets

It is possible to run on local files. This works well when testing or building out your code, but is horrible if you need to run on a large number of files. It is recommended to use this only with a single file. It is, for the most part, a drop-in replacement for the ServiceX backend version.

First, you must install the local variant of func_adl_servicex. If you are using pip, you can do the following:

pip install func_adl_servicex[local]

With that installed, the following will work:

from func_adl_servicex import SXLocalxAOD

dataset_xaod = "my_local_xaod.root"
ds = SXLocalxAOD(dataset_xaod)
data = ds \
    .SelectMany('lambda e: (e.Jets("AntiKt4EMTopoJets"))') \
    .Where('lambda j: (j.pt()/1000)>30') \
    .Select('lambda j: j.pt()') \
    .AsAwkwardArray(["JetPt"]) \
    .value()

print(data['JetPt'])

And replace SXLocalxAOD with SXLocalCMSRun1AOD for using CMS backend (and, of course, update the query).

Development

PR's are welcome! Feel free to add an issue for new features or questions.

The master branch is the most recent commits that both pass all tests and are slated for the next release. Releases are tagged. Modifications to any released versions are made off those tags.

Qastle

This is for people working with the back-ends that run in servicex.

This is the qastle produced for an xAOD dataset:

(call EventDataset 'ServiceXDatasetSource')

(the actual dataset name is passed in the servicex web API call.)

This is the qastle produced for a ROOT flat file:

(call EventDataset 'ServiceXDatasetSource' 'tree_name')

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

func_adl_servicex-2.0b1.tar.gz (8.6 kB view details)

Uploaded Source

Built Distribution

func_adl_servicex-2.0b1-py3-none-any.whl (9.1 kB view details)

Uploaded Python 3

File details

Details for the file func_adl_servicex-2.0b1.tar.gz.

File metadata

  • Download URL: func_adl_servicex-2.0b1.tar.gz
  • Upload date:
  • Size: 8.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.8.0 pkginfo/1.8.2 readme-renderer/34.0 requests/2.27.1 requests-toolbelt/0.9.1 urllib3/1.26.8 tqdm/4.63.0 importlib-metadata/4.11.3 keyring/23.5.0 rfc3986/2.0.0 colorama/0.4.4 CPython/3.8.12

File hashes

Hashes for func_adl_servicex-2.0b1.tar.gz
Algorithm Hash digest
SHA256 506fbd794b1349092050a3402a9a738842f70303189130a8b881983f9a99280c
MD5 89d85914b2c5d1e10a70201408fa458c
BLAKE2b-256 257581923f448332d035c2dc90a7bba0e30cdb9b87f7f4dc2f9df7bc37c256a2

See more details on using hashes here.

Provenance

File details

Details for the file func_adl_servicex-2.0b1-py3-none-any.whl.

File metadata

  • Download URL: func_adl_servicex-2.0b1-py3-none-any.whl
  • Upload date:
  • Size: 9.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.8.0 pkginfo/1.8.2 readme-renderer/34.0 requests/2.27.1 requests-toolbelt/0.9.1 urllib3/1.26.8 tqdm/4.63.0 importlib-metadata/4.11.3 keyring/23.5.0 rfc3986/2.0.0 colorama/0.4.4 CPython/3.8.12

File hashes

Hashes for func_adl_servicex-2.0b1-py3-none-any.whl
Algorithm Hash digest
SHA256 91605bd757412ea111d6012dc35ec0595b7d20bc79164e528203b9cb49baab49
MD5 31ca5e97ac4c9858dd1c17aaac6fe10b
BLAKE2b-256 ddfaafe94ab9939edd8827287eedb6bf4d812f5d763e7998cc80b4d412da7b3c

See more details on using hashes here.

Provenance

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page