Select, weight and analyze complex sample data

These details have not been verified by PyPI

Project links

Project description

Sample Analytics

In large scale surveys, often complex random mechanisms are used to select samples. Estimates derived from such samples must reflect the random mechanism. Samplics is a python package that implements a set of sampling techniques for complex survey designs. These survey sampling techniques are organized into the following four subpackages.

Sampling provides a set of random selection techniques used to draw a sample from a population. It also provides procedures for calculating sample sizes. The sampling subpackage contains:

Sample size calculation and allocation: Wald and Fleiss methods for proportions.
Equal probability of selection: simple random sampling (SRS) and systematic selection (SYS)
Probability proportional to size (PPS): Systematic, Brewer's method, Hanurav-Vijayan method, Murphy's method, and Rao-Sampford's method.

Weighting provides the procedures for adjusting sample weights. More specifically, the weighting subpackage allows the following:

Weight adjustment due to nonresponse
Weight poststratification, calibration and normalization
Weight replication i.e. Bootstrap, BRR, and Jackknife

Estimation provides methods for estimating the parameters of interest with uncertainty measures that are consistent with the sampling design. The estimation subpackage implements the following types of estimation methods:

Taylor-based, also called linearization methods
Replication-based estimation i.e. Boostrap, BRR, and Jackknife
Regression-based e.g. generalized regression (GREG)

Small https://seneweb.com/news/International/france-l-rsquo-esperance-de-vie-et-le-no_n_338526.htmlArea Estimation (SAE). When the sample size is not large enough to produce reliable / stable domain level estimates, SAE techniques can be used to model the output variable of interest to produce domain level estimates. This subpackage provides Area-level and Unit-level SAE methods.

For more details, visit https://samplics.readthedocs.io/en/latest/

Usage

Let's assume that we have a population and we would like to select a sample from it. The goal is to calculate the sample size for an expected proportion of 0.80 with a precision of 0.10.

import samplics
from samplics.sampling import SampleSize

sample_size = SampleSize(parameter = "proportion")
sample_size.calculate(target=0.80, precision=0.10)

Furthermore, the population is located in four natural regions i.e. North, South, East, and West. We could be interested in calculating sample sizes based on region specific requirements e.g. expected proportions, desired precisions and associated design effects.

import samplics
from samplics.sampling import SampleSize

sample_size = SampleSize(parameter="proportion", method="wald", stratification=True)

expected_proportions = {"North": 0.95, "South": 0.70, "East": 0.30, "West": 0.50}
half_ci = {"North": 0.30, "South": 0.10, "East": 0.15, "West": 0.10}
deff = {"North": 1, "South": 1.5, "East": 2.5, "West": 2.0}

sample_size = SampleSize(parameter = "proportion", method="Fleiss", stratification=True)
sample_size.calculate(target=expected_proportions, precision=half_ci, deff=deff)

To select a sample of primary sampling units using PPS method, we can use code similar to:

import samplics
from samplics.sampling import SampleSelection

psu_frame = pd.read_csv("psu_frame.csv")
psu_sample_size = {"East":3, "West": 2, "North": 2, "South": 3}
pps_design = SampleSelection(
   method="pps-sys",
   stratification=True,
   with_replacement=False
   )

frame["psu_prob"] = pps_design.inclusion_probs(
   psu_frame["cluster"],
   psu_sample_size,
   psu_frame["region"],
   psu_frame["number_households_census"]
   )

To adjust the design sample weight for nonresponse, we can use code similar to:

import samplics
from samplics.weighting import SampleWeight

status_mapping = {
   "in": "ineligible",
   "rr": "respondent",
   "nr": "non-respondent",
   "uk":"unknown"
   }

full_sample["nr_weight"] = SampleWeight().adjust(
   samp_weight=full_sample["design_weight"],
   adjust_class=full_sample["region"],
   resp_status=full_sample["response_status"],
   resp_dict=status_mapping
   )

To estimate population parameters, we can use code similar to:

import samplics
from samplics.estimation import TaylorEstimation, ReplicateEstimator

# Taylor-based
zinc_mean_str = TaylorEstimator("mean").estimate(
   y=nhanes2f["zinc"],
   samp_weight=nhanes2f["finalwgt"],
   stratum=nhanes2f["stratid"],
   psu=nhanes2f["psuid"],
   remove_nan=True
)

# Replicate-based
ratio_wgt_hgt = ReplicateEstimator("brr", "ratio").estimate(
   y=nhanes2brr["weight"],
   samp_weight=nhanes2brr["finalwgt"],
   x=nhanes2brr["height"],
   rep_weights=nhanes2brr.loc[:, "brr_1":"brr_32"],
   remove_nan = True
)

To predict small area parameters, we can use code similar to:

import samplics
from samplics.estimation import EblupAreaModel, EblupUnitModel

# Area-level basic method
fh_model_reml = EblupAreaModel(method="REML")
fh_model_reml.fit(
   yhat=yhat, X=X, area=area, intercept=False, error_std=sigma_e, tol=1e-4,
)
fh_model_reml.predict(X=X, area=area, intercept=False)

# Unit-level basic method
eblup_bhf_reml = EblupUnitModel()
eblup_bhf_reml.fit(ys, Xs, areas,)
eblup_bhf_reml.predict(Xmean, areas_list)

Installation

pip install samplics

Python 3.6.1 or newer is required and the main dependencies are numpy, pandas, scpy, and statsmodel.

License

MIT

Contact

created by Mamadou S. Diallo - feel free to contact me!

Project details

These details have not been verified by PyPI

Project links

Release history Release notifications | RSS feed

0.6.0

Mar 10, 2026

0.5.1

Feb 10, 2026

0.5.0

Jan 12, 2026

0.4.55

Aug 22, 2025

0.4.54

Aug 16, 2025

0.4.53

Aug 16, 2025

0.4.52

May 21, 2025

0.4.51

May 17, 2025

0.4.50

May 17, 2025

0.4.49

May 17, 2025

0.4.48

Mar 20, 2025

0.4.47

Mar 19, 2025

0.4.46

Mar 14, 2025

0.4.45

Mar 12, 2025

0.4.44

Feb 26, 2025

0.4.42

Feb 26, 2025

0.4.41

Feb 26, 2025

0.4.40

Feb 25, 2025

0.4.39

Feb 25, 2025

0.4.38

Feb 18, 2025

0.4.37

Feb 14, 2025

0.4.36

Feb 11, 2025

0.4.35

Feb 9, 2025

0.4.34

Jan 31, 2025

0.4.33

Jan 31, 2025

0.4.32

Jan 30, 2025

0.4.31

Jan 6, 2025

0.4.30

Jan 5, 2025

0.4.22

Jul 14, 2024

0.4.21

Jun 19, 2024

0.4.20

Jun 19, 2024

0.4.19

Jun 11, 2024

0.4.18

Jun 11, 2024

0.4.17

Jun 11, 2024

0.4.16

Jun 6, 2024

0.4.15

Jun 6, 2024

0.4.14

May 5, 2024

0.4.13

May 3, 2024

0.4.12

Apr 29, 2024

0.4.11

Dec 10, 2023

0.4.10

Aug 18, 2023

0.4.9

Aug 10, 2023

0.4.8

Jun 3, 2023

0.4.7

Jun 2, 2023

0.4.6

May 2, 2023

0.4.5

Feb 17, 2023

0.4.4

Feb 17, 2023

0.4.3

Feb 17, 2023

0.4.2

Feb 17, 2023

0.4.1

Nov 3, 2022

0.4.0

Nov 1, 2022

0.3.43

Oct 23, 2022

0.3.42

Oct 19, 2022

0.3.41

Sep 25, 2022

0.3.40

Sep 17, 2022

0.3.39

Sep 17, 2022

0.3.38

Jul 12, 2022

0.3.37

Jul 12, 2022

0.3.36

Jun 19, 2022

0.3.35

May 11, 2022

0.3.34

May 11, 2022

0.3.33

May 11, 2022

0.3.32

May 11, 2022

0.3.31

May 11, 2022

0.3.30

May 11, 2022

0.3.29

May 11, 2022

0.3.28

May 11, 2022

0.3.27

May 11, 2022

0.3.26

May 11, 2022

0.3.25

May 11, 2022

0.3.24

May 11, 2022

0.3.23

May 5, 2022

0.3.22

May 5, 2022

0.3.21

May 5, 2022

0.3.20

Apr 30, 2022

0.3.19

Apr 30, 2022

0.3.18

Apr 30, 2022

0.3.17

Apr 29, 2022

0.3.16

Apr 29, 2022

0.3.15

Apr 26, 2022

0.3.14

Apr 26, 2022

0.3.13

Dec 2, 2021

0.3.12

Dec 2, 2021

0.3.11

Oct 29, 2021

0.3.10

Jul 22, 2021

0.3.9

Jul 22, 2021

This version

0.3.8

Apr 28, 2021

0.3.7

Apr 26, 2021

0.3.6

Apr 24, 2021

0.3.5

Apr 19, 2021

0.3.4

Apr 19, 2021

0.3.3

Apr 19, 2021

0.3.2

Feb 22, 2021

0.3.1

Feb 20, 2021

0.3.0

Jan 17, 2021

0.2.6

Oct 18, 2020

0.2.5

Jun 30, 2020

0.2.4

Jun 20, 2020

0.2.3

Jun 17, 2020

0.2.2

Jun 6, 2020

0.2.1

Jun 6, 2020

0.2.0

Jun 6, 2020

0.1.1

May 31, 2020

0.1.0

May 31, 2020

0.0.15

May 24, 2020

0.0.14

May 24, 2020

0.0.13

May 24, 2020

0.0.12

May 24, 2020

0.0.11

May 23, 2020

0.0.10

May 17, 2020

0.0.9

May 16, 2020

0.0.8

May 15, 2020

0.0.7

May 15, 2020

0.0.6

May 15, 2020

0.0.5

May 13, 2020

0.0.4

Jan 29, 2020

0.0.3

Jan 26, 2020

0.0.2

Jan 19, 2020

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

samplics-0.3.8.tar.gz (197.0 kB view details)

Uploaded Apr 28, 2021 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

samplics-0.3.8-py3-none-any.whl (210.1 kB view details)

Uploaded Apr 28, 2021 Python 3

File details

Details for the file samplics-0.3.8.tar.gz.

File metadata

Download URL: samplics-0.3.8.tar.gz
Upload date: Apr 28, 2021
Size: 197.0 kB
Tags: Source
Uploaded using Trusted Publishing? No
Uploaded via: poetry/1.1.4 CPython/3.8.3 Darwin/20.4.0

File hashes

Hashes for samplics-0.3.8.tar.gz
Algorithm	Hash digest
SHA256	`e239cf5e632647db4360fb69b3ef809dc08d23012b472325179bb00534741c76`
MD5	`c15c1a97866ee9f1ee26889045bba4e4`
BLAKE2b-256	`8642d5ea9960dcf466727efe568498794b4d575918ef07c04d80554e91290d8d`

See more details on using hashes here.

File details

Details for the file samplics-0.3.8-py3-none-any.whl.

File metadata

Download URL: samplics-0.3.8-py3-none-any.whl
Upload date: Apr 28, 2021
Size: 210.1 kB
Tags: Python 3
Uploaded using Trusted Publishing? No
Uploaded via: poetry/1.1.4 CPython/3.8.3 Darwin/20.4.0

File hashes

Hashes for samplics-0.3.8-py3-none-any.whl
Algorithm	Hash digest
SHA256	`728711b5b1b2742fad1425a6b018033d759141b858c133c21c9d23be59320e87`
MD5	`40ca7f78436fa99c3510841a1bfba626`
BLAKE2b-256	`85e01109df39ec6b9934ba47eb2548f09a6d47e401ecdcaf98a613831bd4db2b`

See more details on using hashes here.

samplics 0.3.8

Navigation

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Sample Analytics

Usage

Installation

License

Contact

Project details

Verified details

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes