Skip to main content
medicoseacademics@gmail.com
0310-7990649
Research & Biostatistics
Dr Faiza
MBBS (Best Graduate, AIMC Lahore)
FCPS Physiology,
ICMT, CHPE, DHPE (STMU)
MHPE (Riphah International University)
MPH (GC University, Faisalabad)
MBA (Virtual University of Pakistan)
A High Yield Review
Types of Data, Presentation of Data
medicoseacademics@gmail.com
0310-7990649
Research
A process of systematic, scientific data collection, analysis, and interpretation to
find solutions to a problem.
Types:
• Qualitative
• Quantitative
medicoseacademics@gmail.com
0310-7990649
INTRODUCTION TO RESEARCH
Process
Formulating research questions and objectives.
Matching design to objectives.
Defining variables and analysis plans.
Drawing the sample.
Developing tools and collection methods.
Monitoring and data analysis.
Writing the research report.
medicoseacademics@gmail.com
0310-7990649
Topics
• Types of Data
• Data Presentations
• Measures of Central Tendency
• Measures of Dispersion/Variability
• Normal Distribution
• Inferential Statistics
• Precision vs. Accuracy
• Hypothesis Testing Framework
• Choice of Statistical Tests
• Type I and Type II Errors
• Sampling techniques
• Epidemiological study designs
• Epidemiological Measures of
Morbidity and Mortality
• Bias & Confounding Variables
• Inherent Diagnostic Validities
(Sensitivity and Specificityy)
medicoseacademics@gmail.com
0310-7990649
Types of Data
medicoseacademics@gmail.com
0310-7990649
Types of Data
Scale Definition Key Characteristic Examples
Qualitative/
Categorical data
which can’t be
expressed
numerically
Nominal
Data divided into
named categories or
groups.
No implication of order
or ratio.
Male/female,
Smoker/non-smoker.
Ordinal
Data divided into
groups, that can be
placed in a
meaningful order.
No information about
the size of the interval
between ranks.
Student class rankings
(1st, 2nd, 3rd).
Blood pressure (high,
moderate, low)
Quantitative/
Numerical Data
Interval
Data with meaningful
intervals between
measured quantities.
No absolute zero; ratios
of scores are not
meaningful.
Celsius scale
IQ scores.
Ratio
Data with meaningful
intervals and an
absolute zero.
Meaningful ratios exist
(e.g., 120 bpm is twice
as fast as 60 bpm).
Weight, Blood
pressure, Pulse rate
NOIR
medicoseacademics@gmail.com
0310-7990649
Types of Numerical Data
• Discrete Variables:
• Can only take certain values and none in between
• Example
• Number of patients
• Number of syringes used
• Continuous Variables:
• May take any value, typically between certain limits
• Example
• Weight
• Height
• Age
• Blood pressure)
medicoseacademics@gmail.com
0310-7990649
Presentation of Data
medicoseacademics@gmail.com
0310-7990649
Presentation of data
• Data matrix:
• Cases in rows, variables in columns
• Descriptive statistics:
• Summary Data
• Frequency table
• Graphs:
• Bar charts
• Pie charts
• Dot plots
• Histogram (Bell shaped/skewed)
medicoseacademics@gmail.com
0310-7990649
Frequency Distribution
medicoseacademics@gmail.com
0310-7990649
Bar Chart
• Used for Nominal scale data.
• Each rectangle is separated by a
space
• Each bar represents data form
separate categories rather than
continuous groups
medicoseacademics@gmail.com
0310-7990649
Histogram
• Used for Interval or
Ratio scale data.
• Rectangles are
adjacent,
• Present the continuous
nature of the data
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Frequency Polygon
• Created by joining the
midpoints of class
intervals with straight
lines
• Used for Ratio or
Interval data
medicoseacademics@gmail.com
0310-7990649
Bar chart Pie chart
Segmented column chart
Histogram Dot Plot Scatter Plot
Median
Correlation
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Research & Biostatistics
Dr Faiza
MBBS (Best Graduate, AIMC Lahore)
FCPS Physiology,
ICMT, CHPE, DHPE (STMU)
MHPE (Riphah International University)
MPH (GC University, Faisalabad)
MBA (Virtual University of Pakistan)
A High Yield Review
Descriptive Statistics, Normal Distribution
medicoseacademics@gmail.com
0310-7990649
Introduction to
Descriptive Statistics
medicoseacademics@gmail.com
0310-7990649
Descriptive Statistics of Numerical
Variable
medicoseacademics@gmail.com
0310-7990649
Descriptive Statistics of Categorical
Variable
medicoseacademics@gmail.com
0310-7990649
Measures of Central
Tendency
medicoseacademics@gmail.com
0310-7990649
Measures of Central Tendency
Measur
e
Definition Key Characteristics Sensitivity to Outliers
Mode
The observed
value that occurs
with the greatest
frequency.
• Found by simple
inspection;
• it is the highest point on a
frequency polygon.
Distributions can be
bimodal or multimodal.
None: Totally
uninfluenced by small
numbers of extreme
scores.
Uniquely
suited to
describe
distributions
with more than
one peak
Median
The figure that
divides the
distribution in
half when all
scores are listed in
order.
• Also known as the 50th
centile.
• For odd-numbered sets, it
is the middle score;
• for even sets, it is the
average of the two middle
scores.
Low: Insensitive to
extreme scores
because it responds
only to the number of
scores above/below it,
not their actual values.
Best for
skewed data
The arithmetic
• Responds to the exact High: Very sensitive to
medicoseacademics@gmail.com
0310-7990649
Task
• Look at the raw data below representing the number of
final-year MBBS students attending ward classes for one
week.
• No of students (n = 7): 12, 16, 12, 14, 20, 11, 27
• Calculate the Mean, Median, Mode
medicoseacademics@gmail.com
0310-7990649
Solution:
• Step 1: Order the data from lowest to highest
• 11, 12, 12, 14, 16, 20, 27
• Mode: 12 days (The value occurring with the highest frequency).
• Median: 14 days (The middle value of an odd-numbered dataset).
• Mean:
• Mean = 11 + 12 + 12 + 14 + 16 + 20 + 27}/{7} = 112}/7 = 16
*The mean (16) is pulled higher than the median (14) by the outlier 27,
demonstrating a positively skewed distribution.
medicoseacademics@gmail.com
0310-7990649
Measures of Dispersion
medicoseacademics@gmail.com
0310-7990649
Measure Definition
Key
Characteristics
Strengths
(Advantages)
Disadvantages
Range
The difference
between the highest
(maximum) and
lowest (minimum)
scores.
Responds only to
the two extreme
scores in a
distribution.
Easily
determined.
• Highly
distorted by
outliers
• Tends to
increase with
sample size.
Ranges Based
on
Percentiles
Values that divide
ordered data.
• Centiles (100
parts),
• Deciles (10 parts),
• Quartiles (4 parts)
IQR = Q3 –Q1
• States the
percentage of
observations
falling below a
specific score;
• Median is the
50th
percentile.
• Usually
unaffected by
outliers;
• independent of
sample size ;
• highly
appropriate for
skewed data.
• Clumsy to
calculate
• Cannot be
used for small
samples
• Not
algebraically
defined.
medicoseacademics@gmail.com
0310-7990649
Measure Definition
Key
Characteristics
Strengths
(Advantages)
Disadvantages
Variance (σ2
)
The mean of the
squares of all
deviation scores in
the distribution.
Σ(x-x¯)2 / n-1
Quantifies the
amount of
variability or
spread about the
sample mean.
• Uses every
single
observation in
the dataset
• Algebraically
defined.
• Units are the
square of the
raw data,
limiting
intuitive use
• Sensitive to
outliers
Inappropriate
for skewed
data.
Standard
Deviation (σ)
The square root of
the variance.
• Average
distance of an
observation
from the mean
• Same
advantages as
variance
• units match
original data,
easily
• Sensitive to
outliers
• Inappropriate
for highly
skewed data.
medicoseacademics@gmail.com
0310-7990649
A student gets:
• 80 in Physiology
70 in Biochemistry
At first, 80 looks better. But now say:
• Physiology class mean = 75, SD = 5
Biochemistry class mean = 50, SD = 10
For Physiology:
For Biochemistry:
Now explain:
“Although 80 is higher than 70, the student performed better in Biochemistry because
he is 2 SD above the mean there.”
Z-score allows fair comparison between different distributions.
medicoseacademics@gmail.com
0310-7990649
Z-score
• The location of any element expressed in terms of how
many standard deviations it lies above or below the mean.
• Positive Z Score: The element is above the mean.
• Negative Z Score: The element is below the mean.
• Because Z scores are standardized (normalized), they
allow you to compare scores from different normal
distributions that have different means and standard
deviations.
• Z = ±1: Delineates the middle 68% of the distribution.
• Z = ±1.96: Delineates the middle 95% of the distribution.
• Z = ±2.58: Delineates the middle 99% of the distribution.
medicoseacademics@gmail.com
0310-7990649
Task
• Look at the raw data below representing the number of
final-year MBBS students attending ward classes for one
week.
• Dataset (n = 7): 12, 16, 12, 14, 20, 11, 27
• Tasks:
1. Calculate the Range.
2. Calculate the Interquartile Range (IQR).
3. Given that the Sample Mean is 16 and the Sample Standard
Deviation is 3, provide a comprehensive interpretation explaining
how these two metrics together describe student attendance.
medicoseacademics@gmail.com
0310-7990649
• Step 1: Order the data from lowest to highest –
• 11, 12, 12, 14, 16, 20, 27
• 1. Range Calculation:
• Range = Maximum - Minimum value = 27 - 11 = 16
• 2. Interquartile Range (IQR) Calculation:
• Find the Median (Q2): The middle value of the 7 ordered elements is 14 .
• Find the Lower Quartile (Q1):
• This is the median of the lower half of data, Q1 = 12
• Find the Upper Quartile (Q3):
• This is the median of the upper half of data above Q3 = 20
• IQR = Q3 - Q1 = 20 - 12 = 8
• 3. Interpretation Mean = 16 students , Standard Deviation =3 students.
• On average, 16 final-year MBBS students attended the ward classes daily over
the course of the week.
• The Standard Deviation (3) quantifies the variability around that average,
indicating that typical daily attendance deviated above or below that mean of 16
by about 3 students.
medicoseacademics@gmail.com
0310-7990649
Normal Distribution
medicoseacademics@gmail.com
0310-7990649
Symmetrical
Bell-shaped
Normal/Gaussian distribution
medicoseacademics@gmail.com
0310-7990649
Skewed Distribution
medicoseacademics@gmail.com
0310-7990649
Skewed Distribution
medicoseacademics@gmail.com
0310-7990649
95% of the measurements have value
approximately within 2 standard
deviations
(SD) of the mean
medicoseacademics@gmail.com
0310-7990649
Confidence Level:
• The probability that a
calculated interval contains
the true population parameter.
• Common Levels: 90%, 95%,
and 99%.
• To increase precision (making
the interval narrower) at a set
confidence level, the sample
size (n) must be increased.
• 95% CI is essentially
Mean±1.96 Standard Error
medicoseacademics@gmail.com
0310-7990649
Research & Biostatistics
Dr Faiza
MBBS (Best Graduate, AIMC Lahore)
FCPS Physiology,
ICMT, CHPE, DHPE (STMU)
MHPE (Riphah International University)
MPH (GC University, Faisalabad)
MBA (Virtual University of Pakistan)
A High Yield Review
Inferential Statistics, Hypothesis Testing, Types of
Errors, Choice of Correct Statistical Test
medicoseacademics@gmail.com
0310-7990649
Introduction to
Inferential Statistics
medicoseacademics@gmail.com
0310-7990649
Inferential Statistics
• Not Known:
• Population parameters (like Population mean μ and standard deviation σ)
• Known:
• Sample statistics like sample mean ( X ) and standard deviation (S) are known
The task of using a sample to draw conclusions
about a population involves going beyond the
actual information that is available;
↓
It involves inference.
↓
Inferential statistics, therefore, involve using a
statistic to estimate a parameter.
Population (Sampling)
↓
Sample (Testing)
↓
Inference back to Population
medicoseacademics@gmail.com
0310-7990649
Identify Variable Scale - Categorical/numerical
Calculate Sample Statistics Mean/Proportion and
"Spread" SD/Variance
Estimate Sampling Error – To account for random
variation between samples
State Hypotheses – Null and Alternative
Select Test Statistic –Based on data type: chi2
for
proportions, t-test/ANOVA for means.
Compute P-Value – Probability that the results occurred
by chance alone .
Final Decision –p < 0.05, reject Null Hypothesis
(Statistically Significant)
Step 1
Step 2
Step 3
Step 4
Step 5
Step 6
Step 7
medicoseacademics@gmail.com
0310-7990649
Hypotheses
• Hypothesis testing is a structured process to determine if a
result is due to chance or a real effect
• Null Hypothesis (HO):
• The assumption that there is no difference between the groups
being compared.
• Example: A new drug produces no change compared to a placebo.
• Alternative Hypothesis (HA):
• The hypothesis that must be accepted if the null hypothesis is
rejected.
• Example: A new drug has different efficacy compared to a placebo.
medicoseacademics@gmail.com
0310-7990649
Significance Level (alpha)
• The probability level at which the null
hypothesis is rejected.
• By convention, the level is usually set
at p<.05.
If the probability (p) of obtaining a
result by chance is <.05
↓
Null hypothesis is rejected
↓
Result is "statistically significant".
medicoseacademics@gmail.com
0310-7990649
Types of error
Actual Situation
Test Result
Not Rejected Rejected
Null Hypothesis True
False positive /
Type 1 error = Alpha =
0.05
Null Hypothesis False
Type 2 error= beta/
False Negative
Power = 1- beta
Usually, 0.8/80%
Decreasing probability of type 1 error: increases probability of type 2 error
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Power of a Statistical Test
• Power = 1 - beta.
• Acceptable Standard:
• 0.8/80%
• Ways to Increase Power:
1. Increase Sample Size (n)
2. Increase alpha/significance:
• Moving from .01 to .05 makes it easier to reject the null.
3. Increase Effect Size:
• A larger real difference between groups is easier to detect.
4. Decrease Sampling Error:
• Reducing standard deviation (S) lowers the estimated standard error.
medicoseacademics@gmail.com
0310-7990649
Directional Hypothesis
• One tailed t-test is more
powerful than two tailed t
test
• A one-tailed test must
depend on the nature of the
hypothesis being tested and
should therefore be decided
at the outset of the research
• e.g., Drug A has greater
efficacy than the placebo
drug.
medicoseacademics@gmail.com
0310-7990649
Selection of Tests for
Inferential Statistics
medicoseacademics@gmail.com
0310-7990649
Independent / Dependent Variables
• Independent variable:
An attribute or characteristic that influences or affects an
outcome
• Dependent variable:
A dependent variable is an attribute or characteristic that is
dependent on or influenced by the independent variable.
• Alternative names: Outcome/ effect/ criterion/ consequence
variables.
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Inferential
statistic
Independent
Variable
Dependent
Variable
Objective Parametric
test
Non-
parametric
Test
Assess
relationships
Continuous
variable
Continuous
variable
Correlation
between two
continuous
variables
Pearson
correlation
coefficient,
regression
Coefficient
Spearman Rho
Continuous
variable
Categorical Correlation
between
continuous and
binary variables
Discriminant
analysis
Point Biserial
correlation
medicoseacademics@gmail.com
0310-7990649
Inferential
statistic
Independent
Variable
Dependent
Variable
Objective Parametric test Non-parametric
Test
Comparison of
means
Categorical Continuous
variable
comparison of
two means
(independent
groups)
Independent t-
test
Mann-Whitney U
Test (Wilcoxon
Rank-Sum Test)
Categorical Continuous
variable
comparison of
two means
(paired groups)
Paired t-test Wilcoxon
Signed-Rank
Test
Categorical Continuous
variable
comparison of
multiple means
(independent
groups)
Analysis of
variance
(ANOVA,
ANCOVA)
Kruskal-Wallis
Test
Categorical Continuous
variable
comparison of
multiple means
Mixed ANOVA,
Repeated
measure ANOVA
Friedman Test
medicoseacademics@gmail.com
0310-7990649
Inferential
statistic
Independent
Variable
Dependent
Variable
Objective Parametric
test
Non-
parametric
Test
Comparison
of proportions
Categorical Categorical Independent
groups
Chi-square
analysis
Fisher's Exact
test
(for small
samples)
Categorical Categorical paired groups McNemar’s
Test
Exact
McNemar’s
test
(for small
medicoseacademics@gmail.com
0310-7990649
Study
Independent variable /
Dependent variable
Appropriate statistical
test
Is BMI associated with
systolic blood pressure
among adults?
BMI: Continuous
Systolic BP: Continuous
Pearson correlation;
Spearman correlation if
data are non-normal
Is mean haemoglobin
level different between
male and female
students?
Sex: Categorical, two
independent groups
Haemoglobin: Continuous
Independent-samples t-
test;
Mann–Whitney U test if
data are non-normal
Does mean blood
pressure change after
antihypertensive
treatment in the same
patients?
Time: Categorical,
Paired measurements →
Blood pressure:
Continuous
Paired t-test;
Wilcoxon signed-rank test
if data are non-normal
medicoseacademics@gmail.com
0310-7990649
Study
Independent variable /
Dependent variable
Appropriate statistical test
Is mean pain score different
among patients receiving three
different analgesics?
Analgesic group: Categorical,
Three independent groups
Pain score: Continuous
One-way ANOVA;
Kruskal–Wallis test if data are
non-normal
Does mean heart rate change
at baseline, 30 minutes, and 60
minutes post exercise in the
same patients?
Time: Categorical, repeated
measurements
Heart rate: Continuous
Repeated-measures ANOVA;
Friedman test if data are non-
normal
Is the proportion of wound
infection different between
open and laparoscopic surgery
groups?
Type of surgery: Categorical,
independent groups
Wound infection: Categorical
Chi-square test;
Fisher’s exact test when
expected cell counts are small
Does smoking status change
before and after counselling in
the same participants?
Time: Categorical, paired
observations Smoking status:
→
Categorical
McNemar’s test
medicoseacademics@gmail.com
0310-7990649
Research & Biostatistics
Dr Faiza
MBBS (Best Graduate, AIMC Lahore)
FCPS Physiology,
ICMT, CHPE, DHPE (STMU)
MHPE (Riphah International University)
MPH (GC University, Faisalabad)
MBA (Virtual University of Pakistan)
A High Yield Review
Epidemiological Measures, Statistical Decision making
medicoseacademics@gmail.com
0310-7990649
Measure Definition Key Characteristics Examples
Ratio
The division of one
quantity by another
without implying a
specific relationship
between them.
Numerator and
denominator are often
mutually exclusive.
Maternal mortality rate:
number of maternal
deaths per 100,000 live
births
Proportion
A specific type of
ratio where those in
the numerator
must be included
in the denominator.
Expresses a part of a
whole; values usually
range from 0 to 1 or 0%
to 100%.
Fetal deaths out of the
total number of births.
Rate
A proportion that
includes time as an
intrinsic part of the
denominator.
Measures the frequency
of an event in a
population over a
specific period.
Maternal mortality rate:
number of maternal
deaths per 100,000 live
births in one year
medicoseacademics@gmail.com
0310-7990649
Measure Definition Key Characteristics Examples
Prevalence
The proportion of
individuals in a
population who have a
disease at a specific
instant.
A "snapshot" of existing
cases (both old and new).
Includes Point (specific
instant) and Period (total
cases during a window)
types .
P = Number of existing
cases /
Total population
Incidence
The number of new
events or cases that
develop in a population
at risk during a specified
time interval.
Focuses on the arrival of
new cases; acts as a
measure of risk.
Denominator must exclude
those not "at risk" (e.g.,
already immune or lacking
the organ).
Incidence= Number of
new
cases
/
Total
population at
risk
medicoseacademics@gmail.com
0310-7990649
Task: Calculate the Incidence Rate and Prevalence
Rate from the database of a single retirement
community on a specific tracking date.
• Population Matrix:
• Total tracking community size: 500 residents
• Residents with long-standing pre-existing Type 2 Diabetes on Jan
1st: 45 residents
• Healthy, non-diabetic residents on Jan 1st: 455 residents
• New diagnoses of Type 2 Diabetes confirmed between Jan 1st and
Dec 31st: 15 residents
• Task:
• Find the Annual Incidence Rate per 1,000 residents at risk.
• Calculate the Point Prevalence Rate on Dec 31st.
medicoseacademics@gmail.com
0310-7990649
Task: Calculate the Incidence Rate and Prevalence
Rate from the database of a single retirement
community on a specific tracking date.
• Population Matrix:
• Total tracking community size: 500 residents
• Residents with long-standing pre-existing Type 2 Diabetes on Jan 1st: 45 residents
• Healthy, non-diabetic residents on Jan 1st: 455 residents
• New diagnoses of Type 2 Diabetes confirmed between Jan 1st and Dec 31st: 15
residents
Incidence = New cases/population at
risk
= (15/455 )* 1000
= 33 per 1000 residents at
risk
Point Prevalence on Dec 31st = New + old
cases/ Total
population
= (60/500 )* 100
= 12%
medicoseacademics@gmail.com
0310-7990649
Statistics in Medical
Decision Making
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Validity / Accuracy Reliability / Precision / Repeatability
Measures what it is intended to
measure. Close to the true value.
Gives consistent results on repeated
measurements.
Threatened by systematic error or
bias.
e.g. A faulty sphygmomanometer consistently
records blood pressure 10 mmHg higher than the
true value.
Threatened by random error.
e.g., Repeated blood pressure readings vary slightly
because of patient movement or observer variation,
such as 120/80, 124/78, and 118/82 mmHg.
Use gold standard tools, calibration,
clear criteria, representative sampling.
Standardize procedure, train
assessors, repeat measurements, use
checklists.
A calibrated mercury
sphygmomanometer gives the true
blood pressure.
Three readings are close: 120/80,
121/79, 120/81.
medicoseacademics@gmail.com
0310-7990649
Statistical
Parameter
Definition Remarks
Sensitivity
• The ability of a test to detect
disease when it is truly
present
• Very sensitive test is there to
rule out a disease
Specificity
• The ability of a test to detect
the absence of disease when it
is truly absent.
• Highly specific test used for
conforming existence of the
disease
Positive
Predictive
Value
• Probability that a patient with
a positive test actually has the
disease
• As Prevalence , PPV .
↑ ↑
• If a disease is rare, a positive
test is likely a false positive.
Negative
Predictive
Value
• Probability that a patient with
a negative test is actually
disease-free.
• As Prevalence ↓, NPV ↓.
• If a disease is very common, a
negative test is less
trustworthy.
SNOUT (SeNsitive, Negative, rules
OUT)
SPIN (SPecific, Positive, rules
IN)
PPV and NPV change with prevalence, while Sensitivity and
Specificity are properties of the test itself and do not change with
prevalence
medicoseacademics@gmail.com
0310-7990649
Sensitivity Vs Specificity
• HIV Testing Sequence
• Screening: Use a Highly Sensitive test (ELISA) to catch
everyone (rules OUT the healthy).
• Confirmation: Use a Highly Specific test (Western Blot) to
ensure those who tested positive actually have it (rules IN
the sick)
medicoseacademics@gmail.com
0310-7990649
Disease Present
(D+)
Disease Absent
(D-)
Total
Test Positive (T+)
80
(True Positive - TP)
100
(False Positive -
FP)
180
Test Negative (T-)
20
(False Negative -
FN)
800
(True Negative -
TN)
820
Total 100 900 1,000
medicoseacademics@gmail.com
0310-7990649
Disease Present (D+) Disease Absent
(D-)
Total
Test Positive (T+) 80
(True Positive - TP)
100
(False Positive - FP)
180
Positive Predictive Value
=
=80/180=44.4%
Test Negative (T-) 20
(False Negative - FN)
800
(True Negative - TN) 820
Negative Predictive Value
=
=800/820=97.6%
Total 100 900 1,000
Sensitivity=
=80/100=80%
Specificity=
=800/900= 88.9%
medicoseacademics@gmail.com
0310-7990649
Research & Biostatistics
Dr Faiza
MBBS (Best Graduate, AIMC Lahore)
FCPS Physiology,
ICMT, CHPE, DHPE (STMU)
MHPE (Riphah International University)
MPH (GC University, Faisalabad)
MBA (Virtual University of Pakistan)
A High Yield Review
Sample Size & Sampling Techniques, Bias & Confounders
medicoseacademics@gmail.com
0310-7990649
Sampling
medicoseacademics@gmail.com
0310-7990649
Population: All individuals having a
precisely defined set of characteristics
(for example: human, male, age 18–
65, with Stage 3 lung cancer)
Sample: A subset of a defined
population, selected for experimental
study
medicoseacademics@gmail.com
0310-7990649
Sample size
• Means/Proportions:
• Magnitude of designed margin of error (inverse relation)
• A smaller acceptable error requires a larger sample size.
• Confidence level (Direct)
• A higher confidence level requires a larger sample size.
• Variability (Direct)
• Greater standard deviation requires a larger sample size.
• Proportions closer to 50% require a larger sample size.
medicoseacademics@gmail.com
0310-7990649
Sampling Techniques
Probability Sampling Techniques
Everyone has an equal chance to be selected
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Technique Definition Procedure/Technique Example
Simple
Random
Every individual has an
equal probability of
selection.
Assign a unique number
to each member; use a
random number
generator or lottery.
Assigning numbers 1–100
to first graders and
picking 20 via Excel.
Systemati
c
Selection occurs at regular
intervals (nth
individual).
• Calculate interval k =
N/n.
• Randomly pick a
starting point in the
first group, then take
every kth
person.
Choosing every 5th
parent
from a district mailing list
of 1,000 to get a sample of
200.
Stratified
Population is divided into
homogeneous subgroups
(strata) before sampling.
Divide population by a
trait (e.g., gender).
Sample from each stratum
proportionally.
In a group of 3,000 girls
and 6,000 boys, sampling
100 girls and 200 boys to
maintain the 1:2 ratio.
Sampling groups or
Randomly select clusters
Randomly selecting 10
medicoseacademics@gmail.com
0310-7990649
Sampling Techniques
Non-Probability Sampling Techniques
Everyone does not have an equal chance to be
selected
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649Technique Definition Procedure/Technique Example
Consecutive
Including every subject who
meets criteria over a specific
time.
Enroll every eligible patient
who walks into a clinic until
the desired sample size is
met.
Taking every patient with
hypertension seen in a clinic
during the month of
January.
Convenienc
e
Selection based on ease of
access and availability.
Recruit participants who are
"willing and available" in a
haphazard fashion.
A researcher studying
medical students only at the
one institute where they
have permission to visit.
Snowball
Existing participants recruit
future participants from
their networks.
Ask initial participants to
identify or forward the study
to others who meet the
criteria.
Sending a survey to a
superintendent who then
forwards it to their network
of principals.
Purposive
Selection based on the
researcher’s clinical
knowledge or specific intent.
Hand-picking "typical" or
specific cases that represent
the characteristic being
studied.
Specifically interviewing
“distinction holders" to
understand the impact of an
innovative teaching
medicoseacademics@gmail.com
0310-7990649
Task: Match the scenario to the precise
sampling method utilized.
• Sampling Choices: Simple Random, Systematic, Stratified, Cluster, Convenience,
Consecutive.
• Scenarios:
1. A hospital researcher enrolls every single diabetic patient who presents to the
emergency department from June 1st to June 30th.
2. A researcher divides a medical school class roster into male and female
subgroups, then utilizes a computer random number generator to select an
equal proportion of students from each subset.
3. An administrator evaluates clinic wait times by auditing the chart of every 8th
patient checked into the facility.
medicoseacademics@gmail.com
0310-7990649
Task: Match the tactical scenario to the
precise sampling method utilized.
• Scenarios:
1. A hospital researcher enrolls every single diabetic patient who
presents to the emergency department from June 1st to June 30th.
2. A researcher divides a medical school class roster into male and
female subgroups, then utilizes a computer random number
generator to select an equal proportion of students from each
subset.
3. An administrator evaluates clinic wait times by auditing the chart
of every 8th patient checked into the facility.
Consecutive
Sampling
Stratified Random Sampling
Systematic Sampling
medicoseacademics@gmail.com
0310-7990649
Bias and Confounders
medicoseacademics@gmail.com
0310-7990649
Bias
• Any systematic error that results in an incorrect estimate of
the association between the exposure and outcome.
medicoseacademics@gmail.com
0310-7990649
Type of Bias Definition Example How to Reduce / Control
Selection Bias
(Berksonian)
Study participants are not
representative of the target
population
Selecting only hospital
patients for a community
disease study
Random sampling, proper
inclusion criteria, adequate
recruitment methods
Information
Bias
Data are measured or
recorded incorrectly in a
systematic way
Incorrect blood pressure
recording due to faulty
instrument
Standardized measurement
methods, calibrated
instruments, training
Misclassificatio
n Bias
Participants are wrongly
classified regarding
exposure or disease
Diseased patients labeled
as healthy
Clear diagnostic criteria,
validated tools, repeated
assessment
Observer Bias
Researcher consistently
over- or under-reports
findings
Examiner records higher
pain scores for one group
Blinding observers,
structured assessment
forms
Recall Bias
Participants do not
remember past events
accurately
(particularly an issue in
case control studies)
Cancer patients remember
smoking history better
than controls
Use records instead of
memory, shorter recall
period, standardized
interviews
Structured questionnaires,
medicoseacademics@gmail.com
0310-7990649
Type of Bias Definition Example
How to Reduce /
Control
Publication Bias
Positive studies are
more likely to be
published
Negative drug trials
remain unpublished
Trial registration,
systematic reviews,
publication of all
results
Allocation Bias
Unequal assignment
of participants
between groups
Sicker patients placed
into treatment group
Random allocation,
allocation
concealment
Assessment Bias
Outcome assessment
differs between
groups
Knowing treatment
status affects grading
Double blinding,
objective outcome
criteria
Healthy Entrant
Effect
Healthier people are
more likely to
participate
Volunteers healthier
than general
population
Careful sampling,
comparison with
source population
Attrition Bias / Lost-
to-follow-up Bias
Participants drop out
unequally between
groups
More severe cases
leave follow-up
Good follow-up
systems, reminders,
intention-to-treat
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
Confounding/Spurious Variable
• An extraneous variable that "mixes" with the exposure, distorting the
true relationship between the independent and dependent variables.
• It can "hide" a true association or "create" a false (spurious) one.
Studying Smoking and Heart Disease
↓
Alcohol is a confounder, because smokers are more likely to drink
↓
and both factors contribute to heart disease.
medicoseacademics@gmail.com
0310-7990649
Confounder Criteria
• Associated with the Exposure:
• It is related to the independent variable (e.g., smokers are more
likely to consume alcohol).
• Independent Risk Factor:
• It is a risk factor for the disease/outcome on its own (e.g., alcohol
causes heart disease regardless of smoking).
• Not an Intermediate:
• It must not "stand between" the variables in the causal pathway
(unlike a Mediating Variable, which is a necessary step in the
sequence).
medicoseacademics@gmail.com
0310-7990649
Controlling for Confounders
Phase Method Logic
Study Design
Randomization
The "Gold Standard" for RCTs; balances all factors
(known/unknown) equally.
Restriction
Limit the study to one group (e.g., only studying
non-drinkers).
Matching
Ensure both groups have identical levels of the
confounder (e.g., matching by age).
Data Analysis
Stratification
Analyze subgroups separately (e.g., drinkers vs.
non-drinkers).
Multivariable
Regression
Statistically "adjust" for the confounder to see the
"pure" effect of the exposure.
medicoseacademics@gmail.com
0310-7990649
Research & Biostatistics
Dr Faiza
MBBS (Best Graduate, AIMC Lahore)
FCPS Physiology,
ICMT, CHPE, DHPE (STMU)
MHPE (Riphah International University)
MPH (GC University, Faisalabad)
MBA (Virtual University of Pakistan)
A High Yield Review
Epidemiological Study Designs
medicoseacademics@gmail.com
0310-7990649
EPIDEMIOLOGICAL STUDY DESIGNS
Descriptive
Systematic presentation of data to give a clear
picture
Case Report
Case Series
Cross-Sectional
Analytical
Establishing associations or determining risk
factors or causation
Observational:
Cohort,
Case-Control
(Cross-Sectional)
Experimental:
True/Quasi experimental
Causal Comparative
medicoseacademics@gmail.com
0310-7990649
DESCRIPTIVE DESIGN DETAILS
Case Report:
Detailed report of an unusual disease in a single person
(e.g., Thalidomide malformation case).
Case Series:
Description of several unusual cases with similar conditions
No control group is used
Cross-Sectional (Prevalence Study):
Survey of a defined population at a single point in time
Determines the "burden of disease"
Exposure and outcome determined simultaneously.
Temporal relations are hard to establish
medicoseacademics@gmail.com
0310-7990649
OBSERVATIONAL: CASE-CONTROL
Procedure:
• Selection based on outcome
• Cases with disease vs. Controls without
• Assembled to check history of past exposure
Pros
• Inexpensive, quick, and useful for rare diseases or conditions with long development.
Cons:
Selection Bias: Validity depends on whether controls are representative of the general
population's exposure level
Recall Bias: Diseased individuals often recall exposure history differently than healthy
controls
medicoseacademics@gmail.com
0310-7990649
Matching
medicoseacademics@gmail.com
0310-7990649
𝑂𝑑𝑑𝑠 𝑅𝑎𝑡𝑖𝑜=
𝑎 × 𝑑
𝑏 ×𝑐
• Compares odds of outcome between
exposed and unexposed groups
• Case-Control studies use Odds Ratio
(OR), not Relative Risk, because they
cannot measure incidence
medicoseacademics@gmail.com
0310-7990649
OBSERVATIONAL: COHORT STUDIES
Cohort: A group sharing common characteristics followed over time
Classic Example: The Framingham Study (followed ~5000 people over decades for coronary events)
Prospective
Assembles groups in the present, collects baseline, and
follows into the future.
Benefit:
• Establishes temporal association
Measures incidence.
Retrospective
Goes back into history (records) to define risk groups and
follows to the present.
Benefit:
Cheaper and faster than prospective
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
𝑅𝑒𝑙𝑎𝑡𝑖𝑣𝑒 𝑅𝑖𝑠𝑘=
𝑎/(𝑎+𝐶 )
𝑏/(𝑏+ 𝑑)
• The relative risk (RR) indicates the increased (or decreased) risk of disease associated
with exposure to the factor of interest.
• A relative risk of one indicates that the risk is the same in the exposed and unexposed
groups.
• Cohort studies measure Incidence and Relative Risk (RR)
medicoseacademics@gmail.com
0310-7990649
EXPERIMENTAL STUDIES (CLINICAL TRIALS)
Researcher manipulates a situation and measures effects amongst two groups
Proof of Causation: The only design that can actually prove causation
Key Features
Assignment of exposure by the researcher.
Intervention and Control groups.
Random Allocation: Essential for equalizing extraneous variables and guarding against
bias
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649
medicoseacademics@gmail.com
0310-7990649