Skip to content

Commit 78869a4

Browse files
authored
Merge pull request #12 from jihyeon-st/feature/vignettes
Add getting-started vignette
2 parents 9448c3c + 9056f83 commit 78869a4

4 files changed

Lines changed: 224 additions & 1 deletion

File tree

‎.Rbuildignore‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -8,4 +8,6 @@
88
^renv\.lock$
99
^data-raw$
1010

11-
^example_data/
11+
^example_data/
12+
^.*\.Rproj$
13+
^\.Rproj\.user$

‎DESCRIPTION‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -28,3 +28,4 @@ Config/testthat/edition: 3
2828
Roxygen: list(markdown = TRUE)
2929
RoxygenNote: 8.0.0
3030
LazyData: true
31+
VignetteBuilder: knitr

‎vignettes/.gitignore‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,2 @@
1+
*.html
2+
*.R

‎vignettes/clonecensorweighting.Rmd‎

Lines changed: 218 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,218 @@
1+
---
2+
title: "Getting started with clonecensorweighting"
3+
output: rmarkdown::html_vignette
4+
vignette: >
5+
%\VignetteIndexEntry{Getting started with clonecensorweighting}
6+
%\VignetteEngine{knitr::rmarkdown}
7+
%\VignetteEncoding{UTF-8}
8+
---
9+
10+
```{r, include = FALSE}
11+
knitr::opts_chunk$set(
12+
collapse = TRUE,
13+
comment = "#>"
14+
)
15+
```
16+
17+
```{r setup}
18+
library(clonecensorweighting)
19+
```
20+
21+
## Why clone-censor-weighting?
22+
23+
Suppose we want to know whether surgery improves survival in lung cancer
24+
patients. Comparing patients who had surgery against those who did not, using a
25+
naive analysis, can be badly biased. One reason is *immortal time bias*: to
26+
appear in the "surgery" group, a patient has to survive long enough to actually
27+
receive surgery. That guaranteed survival time gets unfairly credited to the
28+
surgery group.
29+
30+
Clone-censor-weighting (CCW) is a technique used in target trial emulation to
31+
address this. The idea has three parts, which also give the method its name:
32+
33+
- **Clone**: each patient is copied into every treatment strategy under
34+
comparison, so that at the start of follow-up everyone is compatible with
35+
every strategy.
36+
- **Censor**: a clone is censored at the moment the patient's observed data
37+
stops being consistent with the strategy that clone was assigned to.
38+
- **Weight**: inverse-probability-of-censoring weights correct for the
39+
artificial censoring introduced in the previous step.
40+
41+
This vignette walks through the pipeline end to end, using a lung cancer
42+
dataset that ships with the package.
43+
44+
## The example data
45+
46+
The package includes `lungcancer`, a dataset of 200 lung cancer patients.
47+
48+
```{r}
49+
data(lungcancer)
50+
head(lungcancer)
51+
```
52+
53+
The columns we rely on are:
54+
55+
- `id`: patient identifier
56+
- `surgery`: whether the patient received surgery (`1`) or not (`0`)
57+
- `timetosurgery`: time from baseline to surgery (`NA` if never operated on)
58+
- `death`: whether the patient died (`1`) or not (`0`) — the outcome event
59+
- `fup_obs`: observed follow-up time
60+
61+
The remaining columns (`sex`, `charlson`, `perf`, `stage`, `emergency`, `age`,
62+
`deprivation`) describe patient characteristics that can be used as covariates
63+
when estimating censoring weights.
64+
65+
We compare two strategies, and give them names we will reuse throughout:
66+
67+
```{r}
68+
arms <- c("Control", "Surgery")
69+
```
70+
71+
The **grace period** is central to this analysis. It is the window during which
72+
a patient is still considered compatible with the "surgery" strategy. A patient
73+
who has surgery within the grace period is consistent with the treated strategy;
74+
one who does not is consistent with the control strategy. Here we use a grace
75+
period of 90 days.
76+
77+
```{r}
78+
grace_period <- 90
79+
```
80+
81+
## Step 1: Clone the data across arms
82+
83+
`clone_arms()` copies the full dataset once per strategy and returns a list with
84+
one data frame per arm.
85+
86+
```{r}
87+
clones <- clone_arms(lungcancer, arms)
88+
89+
names(clones)
90+
```
91+
92+
At this point the two data frames are identical copies of the original data. The
93+
strategy-specific logic is applied in the next steps.
94+
95+
## Step 2: Build the policy and censoring logic
96+
97+
Two helper functions describe how each arm's emulated outcome, follow-up, and
98+
censoring should be derived. They do not touch the data yet; they return the
99+
rules (as expressions) that will be applied in Step 3.
100+
101+
`create_policy_A()` produces the emulated outcome and follow-up under the grace
102+
period policy.
103+
104+
```{r}
105+
policy <- create_policy_A(
106+
arms = arms,
107+
treatment = "surgery",
108+
time_to_treatment = "timetosurgery",
109+
grace_period = grace_period,
110+
outcome = "death",
111+
followup = "fup_obs"
112+
)
113+
```
114+
115+
`create_censoring_logics_A()` produces the censoring indicator and the time at
116+
which that indicator can first be determined.
117+
118+
```{r}
119+
censoring <- create_censoring_logics_A(
120+
arms = arms,
121+
treatment = "surgery",
122+
time_to_treatment = "timetosurgery",
123+
grace_period = grace_period,
124+
followup = "fup_obs"
125+
)
126+
```
127+
128+
Each of these returns a nested list, one entry per arm. The policy rules create
129+
the variables `.outcome` and `.fup`; the censoring rules create `.censoring` and
130+
`.fup_uncensored`.
131+
132+
## Step 3: Apply the logic to the clones
133+
134+
`apply_logics()` evaluates the rules against each arm's data. We combine the
135+
policy and censoring rules for each arm into a single set of new variables.
136+
137+
```{r}
138+
logics <- list(
139+
Control = c(policy$Control, censoring$Control),
140+
Surgery = c(policy$Surgery, censoring$Surgery)
141+
)
142+
143+
emulated <- apply_logics(clones, logics)
144+
145+
lapply(emulated, names)
146+
```
147+
148+
Each arm's data frame now carries the emulated variables (`.outcome`, `.fup`,
149+
`.censoring`, `.fup_uncensored`) alongside the original columns.
150+
151+
## Step 4: Assemble the final long-form data
152+
153+
`create_final_data()` turns each patient into a sequence of time intervals
154+
(a counting-process, or "long", format). Internally it finds every event time,
155+
splits each patient's follow-up at those times, and combines the outcome and
156+
censoring information into one table per arm.
157+
158+
```{r}
159+
final <- create_final_data(
160+
clones = emulated,
161+
clone_followup = ".fup",
162+
clone_outcome = ".outcome",
163+
clone_censoring = ".censoring",
164+
col_ids = "id"
165+
)
166+
167+
head(final$Surgery)
168+
```
169+
170+
Each row is now one patient-interval, bounded by `Tstart` and `Tstop`. Because a
171+
single patient contributes several intervals, each arm has many more rows than
172+
the original 200 patients:
173+
174+
```{r}
175+
lapply(final, dim)
176+
```
177+
178+
This long-form data is the input for the final steps of a CCW analysis:
179+
estimating inverse-probability-of-censoring weights over time and fitting a
180+
weighted survival model to compare the two strategies.
181+
182+
## Reading your own data
183+
184+
The example above uses the bundled `lungcancer` data. To run the same workflow
185+
on your own data, it needs, at a minimum:
186+
187+
- an **identifier** column for each patient
188+
- a binary **treatment** column, coded `0` / `1`
189+
- a **time-to-treatment** column (numeric, `NA` for the untreated)
190+
- a binary **outcome** column, coded `0` / `1`
191+
- a numeric **follow-up time** column
192+
193+
The column *names* are up to you; you pass them to the arguments of the policy
194+
and censoring helpers, as we did above. Any additional columns are carried along
195+
and can serve as covariates.
196+
197+
`read_trial_data()` is a small convenience wrapper for reading such data from a
198+
CSV file into a tibble. It does not impose any particular column structure; it
199+
simply reads the file:
200+
201+
```{r, eval = FALSE}
202+
my_data <- read_trial_data("path/to/your-data.csv")
203+
```
204+
205+
## Summary
206+
207+
In this vignette we:
208+
209+
- motivated clone-censor-weighting with a lung cancer surgery question and the
210+
problem of immortal time bias,
211+
- cloned the data across a control and a surgery arm with `clone_arms()`,
212+
- described the grace-period policy and censoring rules with `create_policy_A()`
213+
and `create_censoring_logics_A()`,
214+
- applied those rules with `apply_logics()`,
215+
- and assembled a long-form analysis dataset with `create_final_data()`.
216+
217+
The remaining steps of a full CCW analysis — estimating censoring weights and
218+
fitting a weighted outcome model — build directly on this long-form dataset.

0 commit comments

Comments
 (0)