576 lines
36 KiB
TeX
576 lines
36 KiB
TeX
%! Author = alex
|
||
%! Date = 3/7/25
|
||
|
||
|
||
\section{Background}\label{sec:background}
|
||
|
||
\subsection{Physiological Background}\label{subsec:physiological_background}
|
||
|
||
\subsubsection{Menstrual Cycle}\label{subsec:menstrual_cycle}
|
||
The menstrual cycle consists of physiological changes preparing the female body for potential pregnancy,
|
||
typically spanning around 28 days but varying considerably among individuals.
|
||
It includes two main phases: the follicular phase, beginning with menstruation, and the luteal phase, following ovulation.
|
||
|
||
During the follicular phase, ovarian follicles mature under the influence of rising estradiol levels, thickening the uterine lining (endometrium).
|
||
Around mid-cycle, a surge of luteinizing hormone (LH) and follicle-stimulating hormone (FSH), triggered by peak estradiol,
|
||
induces ovulation—the release of a mature egg into the fallopian tube.
|
||
|
||
After ovulation, the luteal phase begins.
|
||
Progesterone increases substantially, maintaining endometrial thickness for potential embryo implantation.
|
||
In parallel, a subtle rise in body temperature (~0.5°C) occurs due to progesterone elevation.
|
||
If fertilization does not happen, progesterone and temperature decline back to baseline levels, resulting in menstruation and initiating a new cycle.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.6\textwidth]{resources/figures/background/background_menstrual_cycle_physiology}
|
||
\caption{Physiological changes during the menstrual cycle~\cite{pedroso_menstrual_2022}.
|
||
The figure shows the relationship between hormonal fluctuations, physiological changes, and body temperature, with a clear temperature rise around ovulation.}
|
||
|
||
\label{fig:background_menstrual_cycle_physiology}
|
||
\end{figure}
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_labeled_cycle}
|
||
\caption{Body core temperature curve across a menstrual cycle. The red line represents a locally smoothed temperature trend.
|
||
Menstruation, ovulation, and the fertile window are indicated in red, blue, and green, respectively.}
|
||
\label{fig:background_labeled_cycle}
|
||
\end{figure}
|
||
|
||
Figure~\ref{fig:background_menstrual_cycle_physiology} illustrates these physiological changes, highlighting hormonal fluctuations and temperature shifts around ovulation.
|
||
The menstrual cycle length varies significantly, influenced by factors such as stress, age, diet, and exercise~\cite{silberstein_physiology_2000}.
|
||
|
||
Figure~\ref{fig:background_labeled_cycle} shows the temperature curve over the course of a menstrual cycle with the
|
||
menstruation, fertile window and ovulation marked.
|
||
It starts with a menstruation and ends just before the next menstruation.
|
||
The follicular phase starts at the beginning and goes on until the ovulation.
|
||
The luteal phase begins at the ovulation and continues until the next menstruation.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_anovulatory_cycle}
|
||
\caption{Example of a cycle without an ovulation and the resulting absence of a temperature rise}
|
||
\label{fig:background_anovulation}
|
||
\end{figure}
|
||
|
||
|
||
While many cycles exhibit a characteristic biphasic pattern, deviations from this norm are common.
|
||
Some remain monophasic, which might be an indication for an anovulatory cycle, which is a menstrual cycle, where no ovulation occurs.
|
||
Anovulation can have various causes, including hormonal imbalances, stress, or underlying health conditions~\cite{rosenfield_adolescent_2013}.
|
||
Monophasic cycles with a confirmed ovulation event have been observed, so there seems to be no clear indication that it is a direct cause of anovulation~\cite{moghissi_accuracy_1976}.
|
||
Thus, distinguishing between ovulatory and anovulatory cycles is challenging, as the only definitive confirmation of
|
||
successful ovulation in a clinical sense is a pregnancy.
|
||
Even ultrasound imaging can only confirm that an egg was released from its follicle—not whether it was successfully implanted or fertilized.
|
||
Figure~\ref{fig:background_anovulation} shows a cycle with a monophasic temperature pattern.
|
||
|
||
To illustrate the diversity of real-world menstrual cycles, Figures~\ref{fig:background_long_cycle} and~\ref{fig:background_short_cycle}
|
||
show examples of cycles that are significantly longer or shorter than a normative 28-day cycle.
|
||
\citeauthor{bull_real-world_2019} have done an extensive study on cycle variability,
|
||
highlighting that women frequently deviate from the normative cycle, especially with age\cite{bull_real-world_2019}.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_long_cycle}
|
||
\caption{Example of a long cycle with a length of 111 days}
|
||
\label{fig:background_long_cycle}
|
||
\end{figure}
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_short_cycle}
|
||
\caption{Example of a short cycle with a length of 22 days}
|
||
\label{fig:background_short_cycle}
|
||
\end{figure}
|
||
|
||
These irregularities appear not only on a per-cycle basis, but also across time within the same individual.
|
||
Figures~\ref{fig:background_irregular_cycles} and~\ref{fig:background_regular_cycles}
|
||
show examples of a woman with an irregular and a regular menstrual cycle pattern, respectively.
|
||
The irregular example highlights how multiple parameters can vary between individuals:
|
||
cycle length, timing of ovulation, temperature shift magnitude between phases, and intra-phase temperature variability.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_irregular_cycle_example}
|
||
\caption{Example of a woman with irregular menstrual rhythm. The dashed vertical lines indicate the ends of each cycle.}
|
||
\label{fig:background_irregular_cycles}
|
||
\end{figure}
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_regular_cycle_example}
|
||
\caption{Example of a woman with regular menstrual rhythm. The dashed vertical lines indicate the ends of each cycle.}
|
||
\label{fig:background_regular_cycles}
|
||
\end{figure}
|
||
|
||
\subsubsection{Fertility Prediction}\label{subsubsec:fertility_prediction}
|
||
Fertility varies throughout the menstrual cycle, centered around the ovulation event.
|
||
An egg remains viable for about 24 hours post-ovulation, while sperm can survive up to 6 days in a woman's reproductive tract,
|
||
thus the fertile period extends to approximately five days prior to ovulation~\cite{dunson_day-specific_1999}.
|
||
Consequently, the whole fertile window generally spans six days: five days preceding ovulation and one day after.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.7\textwidth]{resources/figures/background/background_pregnancy_chance_over_time}
|
||
\caption{Chance of fertilization depending on the day of the menstrual cycle.
|
||
The highest chance is around one day before ovulation~\cite{dunson_day-specific_1999}.}
|
||
\label{fig:background_pregnancy_chance}
|
||
\end{figure}
|
||
|
||
Figure~\ref{fig:background_pregnancy_chance} demonstrates the probability of fertilization peaking one day before ovulation, emphasizing the critical timing for fertility prediction.
|
||
|
||
Fertility prediction fundamentally depends on accurate ovulation timing.
|
||
However, since the goal is to identify the fertile window before ovulation occurs, detection must be early and precise.
|
||
For individuals trying to conceive or avoid pregnancy, knowing the window of fertility is more actionable than identifying the ovulation event itself.
|
||
|
||
\subsubsection{Practical Use Cases}\label{subsubsec:practical_use_cases}
|
||
In this study, we will focus on \emph{natural family planning} (NFP), which includes preventing and achieving pregnancy.
|
||
Individuals aiming to avoid pregnancy identify fertile days to abstain from intercourse,
|
||
whereas those seeking pregnancy aim to focus intercourse around days with the highest fertility probability.
|
||
|
||
Both use cases revolve around accurately predicting ovulation.
|
||
However, the implications of prediction errors differ significantly.
|
||
A false-positive prediction indicates high fertility despite actual fertility being low or nonexistent,
|
||
whereas a false-negative prediction implies low fertility when fertility is actually high.
|
||
|
||
For women aiming to avoid pregnancy, minimizing false-negative predictions is crucial due to the risk of unintended pregnancy.
|
||
Although false-positives may lead to unnecessary abstinence, this outcome is generally considered less severe.
|
||
Consequently, prediction algorithms should be conservative, erring on the side of higher fertility estimates to prioritize safety.
|
||
|
||
For women aiming to conceive, incorrect predictions are less severe, as they might, at worst, result in misdirected
|
||
pregnancy effort, causing frustration or delays, but no potential dangers to the health of a woman.
|
||
Therefore, algorithms for this group don't need to be as conservative.
|
||
It remains to be seen, where the middle ground lies and how different algorithms perform for different use cases.
|
||
|
||
\subsubsection{Physiological Signs of Ovulation}\label{subsubsec:physiological_signs}
|
||
Several physiological signs correlate with ovulation and can be used for prediction.
|
||
As shown in Figure~\ref{fig:background_menstrual_cycle_physiology}, these include hormonal fluctuations (LH and FSH surges) and
|
||
changes in body temperature.
|
||
Additionally, variations in cervical mucus consistency, salivary ferning patterns,
|
||
and electrical resistance of the skin and vaginal mucosa have been observed~\cite{silberstein_physiology_2000}.
|
||
|
||
Among these, ultrasonography provides the most accurate confirmation of ovulation by detecting follicular changes,
|
||
but it is costly and requires specialized equipment.
|
||
Hormone measurements in urine and blood are widely used and available in at-home test kits,
|
||
but they require frequent testing.
|
||
Temperature-based methods, particularly body temperature tracking, offer a non-invasive alternative by
|
||
detecting the slight temperature rise that follows ovulation.
|
||
Advances in wearable technology have further enabled continuous and automated temperature monitoring,
|
||
improving accessibility and usability~\cite{alexander_fertilitatsmonitoring_2014, luo_detection_2020, yu_tracking_2022}.
|
||
|
||
Among the available methods, temperature-based monitoring, especially when automated and continuous, offers a promising avenue for large-scale cycle analysis.
|
||
|
||
\subsection{Data Source and Characteristics}\label{subsec:data_background}
|
||
This study is based on a dataset collected from users of the \emph{OvulaRing}~\cite{noauthor_ovularing_nodate},
|
||
an intravaginal wearable sensor developed by VivoSensMedical GmbH, located in Leipzig, Germany~\cite{noauthor_vivosens_nodate}.
|
||
VivoSensMedical is supported by the European Regional Development Fund (EFRE), Saxon State Ministry for Economic Affairs,
|
||
Labour and Transport (SMWA) and the State Ministry of Science and Fine Arts (SMWK) and the City of Leipzig.
|
||
OvulaRing has been financed by the startup grant `EXIST Gründerstipendium` from the Federal Ministry of Economics and Technology.
|
||
The device continuously records intravaginal core body temperature at 5-minute intervals.
|
||
The sensor itself measures approximately 1\,cm $\times$ 1\,cm $\times$ 2\,cm and is embedded in a silicone ring with a diameter of 5\,cm for ease of use.
|
||
It pairs with a smartphone via Bluetooth to synchronize and upload recorded data to a secure database.
|
||
Figure~\ref{fig:background_ovularing} shows an image of the ring attached to its silicone ring.
|
||
|
||
The product has been on the market for over a decade, resulting in an extensive longitudinal dataset of menstrual cycles.
|
||
Cycle boundaries are defined by self-reported menstruation, which users manually log in the accompanying app to mark the beginning of each cycle.
|
||
|
||
Thanks to a battery life of at least six months, the device supports continuous monitoring of long and irregular cycles, enabling the capture of highly variable menstrual patterns.
|
||
|
||
A known limitation of manual cycle annotations is the potential for misalignment.
|
||
Intermediate bleeding events unrelated to menstruation (e.g., ovulatory spotting or irregular shedding) or missing menstruation entries can lead to ambiguous cycle definitions.
|
||
Therefore, all user-entered cycle starts undergo manual review to reduce annotation errors.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.3\textwidth]{resources/figures/background/ovularing}
|
||
\caption{The OvulaRing sensor attached to its silicone ring~\cite{noauthor_ringpng_nodate}.}
|
||
\label{fig:background_ovularing}
|
||
\end{figure}
|
||
|
||
In addition to temperature measurements, the database includes contextual metadata such as age, height, weight, and optional user-entered markers.
|
||
These markers provide further physiological context and may include information about intermediate bleeding, sexual intercourse, or positive pregnancy tests.
|
||
|
||
At the time of writing, the dataset contains approximately 65{,}000 annotated cycles,
|
||
comprising more than 350 million individual temperature measurements~\footnote{This is the largest dataset of continuous body core temperature data used in any study so far.}.
|
||
|
||
\subsubsection{Dataset Summary}
|
||
For the present study, the dataset was reduced to approximately 40{,}000 cycles after filtering out entries that were incomplete,
|
||
contained hardware-related anomalies, or fell outside a reasonable cycle length range.
|
||
|
||
Cycles shorter than 10 days typically result from incorrect cycle start entries or premature termination of temperature recordings.
|
||
Long cycles, longer than 150 days, are often due to data entry errors or pregnancy-related recordings,
|
||
where the sensor was worn continuously throughout gestation—sometimes producing sequences up to nine months long.
|
||
|
||
While such cases may still contain useful information, they were excluded from this analysis to avoid complications in preprocessing and labeling.
|
||
In most instances, only a small portion of these extended cycles contributes meaningfully to the study objectives.
|
||
The cutoff values for cycle length are 10 and 150 days, respectively.
|
||
|
||
The cleaned dataset has 40{,}266 menstrual cycles from 6{,}245 users.
|
||
The median number of cycles per user is 4 (IQR: 2--8) and the median cycle length is 28 days (IQR: 26--32).
|
||
|
||
The average data density—defined as the fraction of available measurements out of the theoretical maximum of 288 measurements per day—is 0.90.
|
||
This corresponds to an average data availability of 90\% per cycle, with an average loss of 10\%.
|
||
|
||
37934 cycles (94\%) were classified as biphasic and 2333 (6\%) as monophasic.
|
||
|
||
Users had a median age of 32 years (IQR: 29–36), median weight of 65 kg (IQR: 58–77), and median height of 168 cm (IQR: 163–172).
|
||
|
||
\subsubsection{Irregularities and Confounding Factors}
|
||
Despite careful data collection, real-world measurements are subject to physiological and behavioral noise.
|
||
Especially core body temperature is influenced by various factors unrelated to the menstrual cycle.
|
||
Illnesses—especially those involving fever—can significantly affect temperature patterns.
|
||
This poses a challenge for any analysis relying on temperature data, as one of the key physiological indicators of ovulation is a post-ovulatory temperature rise (see Section~\ref{subsubsec:physiological_signs}).
|
||
|
||
Figure~\ref{fig:background_fever_cycle} shows an example of a cycle where an illness caused a marked increase in temperature.
|
||
This event is particularly problematic because the fever-induced rise occurs just before the expected ovulatory shift, potentially confounding fertility detection.
|
||
Distinguishing illness-related changes from cycle-related ones requires models that are sensitive to context and robust to outliers.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_fever_cycle}
|
||
\caption{Example of a cycle affected by illness, showing a fever-induced temperature rise shortly before the expected ovulatory shift.}
|
||
\label{fig:background_fever_cycle}
|
||
\end{figure}
|
||
|
||
Another common irregularity arises when users temporarily remove the sensor during menstruation, typically for hygiene reasons—even though the device is safe to wear continuously.
|
||
This behavior frequently results in missing data at the start of each cycle.
|
||
Depending on the total cycle length, this gap can represent a significant portion of the cycle's data.
|
||
|
||
Figure~\ref{fig:background_menstruation_data_gap} shows an example cycle with a data gap during menstruation.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{resources/figures/background/background_menstruation_data_gap}
|
||
\caption{Example of a cycle with a data gap at the beginning, caused by sensor removal during menstruation.}
|
||
\label{fig:background_menstruation_data_gap}
|
||
\end{figure}
|
||
|
||
\subsubsection{Privacy}
|
||
The dataset used in this study contains sensitive personal health information and is handled with strict privacy safeguards.
|
||
All data is pseudonymized and processed exclusively on encrypted devices, ensuring that no identifiable information can be traced back to individual users.
|
||
VivoSensMedical does not share user data with third parties; the data is used solely for internal research and product improvement efforts that directly benefit users at no additional cost.
|
||
|
||
\subsection{Technical Background}\label{subsec:technological_background}
|
||
With a large, high-resolution dataset of longitudinal temperature measurements and associated metadata available,
|
||
the next challenge lies in how to model such sequential data effectively.
|
||
Accurate ovulation prediction requires algorithms that can handle temporal dependencies,
|
||
irregularities, and physiological variability across users.
|
||
To this end, we turn to machine learning techniques designed for time series analysis,
|
||
beginning with foundational concepts and progressing to modern neural architectures.
|
||
|
||
\subsubsection{Time Series Analysis}\label{subsubsec:time_series_analysis}
|
||
Temperature data recorded by the OvulaRing forms a high-resolution time series, where each measurement carries temporal context.
|
||
Time series analysis is essential to uncover meaningful patterns and predict future physiological states from such sequential data.
|
||
In machine learning, this often involves models that can learn temporal dependencies—most notably
|
||
Recurrent Neural Networks (RNNs) and the more recent Transformer architecture.
|
||
|
||
\subsubsection{RNN and LSTM Networks}\label{subsubsec:lstm_networks}
|
||
Recurrent Neural Networks (RNNs) are extensions of classical neural networks that incorporate cyclic connections between neurons.
|
||
These recurrent connections allow the network to retain information from previous inputs by feeding the hidden state
|
||
from a prior time step into the current one, enabling a form of temporal memory.
|
||
They process input \emph{sequentially}, maintaining this hidden state over time.
|
||
|
||
In practice, this means that input data is processed sequentially, one step at a time.
|
||
At each step \(t\), the input \(x_t\) is combined with the previous hidden state \(h_{t-1}\) to produce a new
|
||
hidden state \(h_t\), which contributes to the output \(o_t\).
|
||
This enables the network to learn temporal dependencies in sequential data~\cite{medsker_recurrent_1999}.
|
||
|
||
However, this approach has the downside that the model cannot explicitly control how it remembers or forgets information at each step,
|
||
limiting its ability to manage long-term dependencies.
|
||
During training via backpropagation, the weights of a neural network are updated based on the partial derivatives of the loss function.
|
||
The more propagations required (e.g., in deeper networks), the more multiplications are needed to compute gradients for earlier weights.
|
||
|
||
RNNs are typically trained using \textit{backpropagation through time} (BPTT), in which gradients are propagated through many time steps,
|
||
leading to numerical instability.
|
||
If the gradients shrink exponentially, the model suffers from the \emph{vanishing gradient} problem;
|
||
if they grow exponentially, it results in \emph{exploding gradients}~\cite{hochreiter_vanishing_1998}.
|
||
In both cases, learning is significantly impaired.
|
||
|
||
Exploding gradients can often be mitigated using techniques such as \emph{gradient clipping},
|
||
where the magnitude of the gradient is capped---typically within a range of \([-1, 1]\)---to stabilize training.
|
||
|
||
Figure~\ref{fig:rnn_unfolded} illustrates the unfolded structure of an RNN across three time steps.
|
||
This technique, known as \emph{unfolding}, clarifies how sequential inputs update the hidden state and generate
|
||
outputs at each time step.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.8\textwidth]{resources/figures/background/recurrent_neural_network_unfold}
|
||
\caption{Schematic diagram of the unfolded structure of a recurrent neural network~\cite{fdeloche_english_2017}}
|
||
\label{fig:rnn_unfolded}
|
||
\end{figure}
|
||
|
||
To address the vanishing gradient problem and enable better long-term memory,
|
||
\emph{Long Short-Term Memory} (LSTM) networks were introduced by \citeauthor{hochreiter_long_1997}~\cite{hochreiter_long_1997}.
|
||
LSTMs extend the RNN architecture by incorporating a memory cell and a series of gates that regulate the flow of
|
||
information: the \emph{forget gate}, the \emph{input gate}, and the \emph{output gate}.
|
||
|
||
\begin{itemize}
|
||
\item The \emph{forget gate} determines which information from the previous cell state should be discarded.
|
||
\item The \emph{input gate} controls what new information is added to the cell state.
|
||
\item The \emph{output gate} selects relevant parts of the current cell state to produce the output and the next hidden state.
|
||
\end{itemize}
|
||
|
||
Each gate employs a sigmoid activation function to regulate the flow of information,
|
||
allowing LSTMs to preserve and update memory over long sequences.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.6\textwidth]{resources/figures/background/background_lstm_structure}
|
||
\caption{Architecture of a memory cell \( c_j \) and its gate units \( \text{in}_j \) and \( \text{out}_j \).
|
||
At the center, the self-recurrent linear unit (with weight 1.0) maintains the internal state~\cite{hochreiter_long_1997}.
|
||
See text for more details.}
|
||
\label{fig:lstm_architecture}
|
||
\end{figure}
|
||
|
||
The LSTM unit maintains an internal state \( s_{c_j}(t) \) for each cell $c_j$ that evolves over time to capture long-term dependencies.
|
||
Figure~\ref{fig:lstm_architecture} shows the structure of a memory cell along with its input and output gates.
|
||
The \emph{input gate} activation \( y_j^{\text{in}}(t) \) controls the influence of the new candidate input \( g(\text{net}_{c_j}(t)) \) on the state update (Eq.~\eqref{eq:state_update}):
|
||
|
||
\begin{equation}
|
||
s_{c_j}(t) = s_{c_j}(t - 1) + y_j^{\text{in}}(t) \cdot g(\text{net}_{c_j}(t))
|
||
\label{eq:state_update}
|
||
\end{equation}
|
||
|
||
The \emph{output gate} \( y_j^{\text{out}}(t) \), computed from \( \text{net}_{\text{out}_j}(t) \), determines how much of the transformed internal state is passed on as visible output (Eq.~\eqref{eq:cell_output}):
|
||
|
||
\begin{equation}
|
||
y_j^{c}(t) = y_j^{\text{out}}(t) \cdot h(s_{c_j}(t))
|
||
\label{eq:cell_output}
|
||
\end{equation}
|
||
|
||
\(g\) and \(h\) are differentiable functions, here \emph{tanh}.
|
||
The gate activations themselves are derived from net inputs formed as weighted sums of previous outputs \( y^u(t-1) \), as given by:
|
||
|
||
\begin{align}
|
||
\text{net}_{\text{out}_j}(t) &= \sum_u w_{\text{out}_j,u} \, y^u(t - 1) \label{eq:net_out} \\
|
||
\text{net}_{\text{in}_j}(t) &= \sum_u w_{\text{in}_j,u} \, y^u(t - 1) \label{eq:net_in} \\
|
||
\text{net}_{c_j}(t) &= \sum_u w_{c_j,u} \, y^u(t - 1) \label{eq:net_c}
|
||
\end{align}
|
||
|
||
During training, the weights associated with the inputs \( x_t \) and previous outputs \( h_{t-1} \),
|
||
as well as the biases for each gate in a cell, are learned through backpropagation.
|
||
To build more expressive models, memory cells can be stacked in multiple layers, and their outputs concatenated or passed sequentially to higher layers.
|
||
|
||
LSTMs are widely used in biomedical applications due to their capacity to handle sequences of variable length and complexity.
|
||
In the context of fertility prediction, where hormonal patterns exhibit periodicity but also irregularity,
|
||
LSTMs are well-suited to learn relevant time-dependent signals from sequential physiological measurements.
|
||
|
||
While powerful, LSTMs can be computationally intensive and sensitive to hyperparameter tuning.
|
||
Therefore, they are often compared with alternative architectures,
|
||
including simpler feedforward networks and more recent attention-based models,
|
||
to evaluate trade-offs in performance, interpretability, and computational cost.
|
||
|
||
Given their ability to learn from sequences with noisy periodic structure,
|
||
LSTMs offer a natural choice for modeling hormonal and temperature fluctuations across menstrual cycles.
|
||
|
||
The next section introduces the \emph{Transformer} architecture, a more recent alternative that forgoes
|
||
recurrence in favor of attention mechanisms.
|
||
|
||
\subsubsection{Transformer Models}\label{subsubsec:transformer_models}
|
||
|
||
Transformer models are a class of neural architectures that use \emph{self-attention}
|
||
to model dependencies in sequential data without relying on recurrence~\cite{vaswani_attention_2017}.
|
||
Unlike recurrent neural networks (RNNs), Transformers process input sequences \emph{in parallel},
|
||
allowing them to model relationships between any pair of input tokens or timesteps directly.
|
||
This design mitigates the limitations of recurrent models, such as long-term memory constraints
|
||
and vanishing gradients.
|
||
|
||
Originally introduced for machine translation, Transformers have proven broadly applicable to
|
||
various sequence modeling tasks due to their flexibility, scalability, and strong performance
|
||
on complex temporal patterns.
|
||
|
||
At the core of the Transformer is the attention mechanism, which enables the model to compute
|
||
context-aware representations by weighing the importance of different input positions for each output.
|
||
This is achieved through \emph{scaled dot-product attention}, where queries, keys, and values are
|
||
linearly projected from the input and used to compute attention scores.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.4\textwidth]{resources/figures/background/background_transformer_architecture}
|
||
\caption{The Transformer - architecture for an encoder-decoder model~\cite{vaswani_attention_2017}}
|
||
\label{fig:background_transformer_architecture}
|
||
\end{figure}
|
||
|
||
|
||
Figure~\ref{fig:background_transformer_architecture} illustrates the original encoder-decoder model introduced by~\citeauthor{vaswani_attention_2017}.
|
||
|
||
The Transformer architecture consists of two components: an \emph{Encoder} and a \emph{Decoder}.
|
||
|
||
\paragraph{Encoder:}
|
||
|
||
The encoder is responsible for encoding the input into a contextualized representation.
|
||
In the case of machine translation, this input would be a sentence in the source language.
|
||
|
||
|
||
The input tokens are first mapped to dense continuous vector representations (embeddings).
|
||
Since the attention mechanism permutation-invariant---that is, it does not inherently encode the order of tokens in the sequence---
|
||
\emph{positional encodings} are added to the token embeddings to provide information about the token positions in the sequence.
|
||
|
||
|
||
Without positional encoding, repeated tokens such as `The` would be indistinguishable
|
||
to the model regardless of their location, even if they play different syntactic or semantic roles.
|
||
Positional encodings, often based on sinusoidal functions, inject a unique position-dependent signal
|
||
into each token, enabling the model to distinguish between identical tokens in different positions.
|
||
|
||
In this work, we use sine and cosine functions of different frequencies:
|
||
\begin{align}
|
||
PE_{\text{pos}, 2i} &= \sin\left(\frac{\text{pos}}{10000^{\frac{2i}{d_{\text{model}}}}}\right), \\
|
||
PE_{\text{pos}, 2i+1} &= \cos\left(\frac{\text{pos}}{10000^{\frac{2i}{d_{\text{model}}}}}\right)
|
||
\end{align}
|
||
where \(i\) is the dimension of the input and \(pos\) is the position in the sequence,
|
||
as in the original paper~\cite{vaswani_attention_2017}.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.8\textwidth]{resources/figures/background/background_attention_mechanism}
|
||
\caption{(left) Scaled Dot-Product Attention.
|
||
(right) Multi-Head Attention consisting of several attention layers running in parallel~\cite{vaswani_attention_2017}.}
|
||
\label{fig:background_attention_mechanism}
|
||
\end{figure}
|
||
|
||
Inside each encoder block, \emph{Multi-Head Attention} is applied to the input sequence.
|
||
This mechanism allows the model to jointly attend to information from different representation subspaces at different positions.
|
||
The process begins by linearly projecting the input sequence $X \in \mathbb{R}^{n \times d_{\text{model}}}$ into three distinct matrices:
|
||
|
||
\begin{align}
|
||
Q &= XW^Q \\
|
||
K &= XW^K \\
|
||
V &= XW^V
|
||
\end{align}
|
||
|
||
where $W^Q$, $W^K$, and $W^V$ are learned projection matrices.
|
||
Each row of $X$ corresponds to the embedding of one token, and the projections produce the queries ($Q$), keys ($K$), and values ($V$) for each token.
|
||
|
||
The core operation of attention is then computed as:
|
||
|
||
\begin{equation}
|
||
\text{Attention}(Q, K, V) = \text{softmax}\left( \frac{QK^\top}{\sqrt{d_k}} \right) V
|
||
\end{equation}
|
||
|
||
The dot product $QK^\top$ measures the pairwise similarity between queries and keys, scaled by the dimensionality $d_k$ to control the magnitude.
|
||
The softmax normalizes the scores across each row, producing an attention distribution over all input positions.
|
||
|
||
Multi-head attention runs this mechanism in parallel across $h$ different heads.
|
||
Each head uses different learned projections:
|
||
|
||
\begin{equation}
|
||
\text{head}_i = \text{Attention}(XW_i^Q, XW_i^K, XW_i^V)
|
||
\end{equation}
|
||
|
||
The resulting vectors from all heads are concatenated and projected:
|
||
|
||
\begin{equation}
|
||
\text{MultiHead}(X) = \text{Concat}(\text{head}_1, \ldots, \text{head}_h)W^O
|
||
\end{equation}
|
||
|
||
This output is then passed through a position-wise feed-forward network (FFN), applied independently to each token:
|
||
|
||
\begin{equation}
|
||
\text{FFN}(x) = \max(0, xW_1 + b_1)W_2 + b_2
|
||
\end{equation}
|
||
|
||
To stabilize training, the Transformer uses residual connections and layer normalization:
|
||
|
||
\begin{align}
|
||
x_1 &= \text{LayerNorm}(X + \text{MultiHead}(X)) \\
|
||
x_2 &= \text{LayerNorm}(x_1 + \text{FFN}(x_1))
|
||
\end{align}
|
||
|
||
In summary, each encoder block consists of:
|
||
\begin{itemize}
|
||
\item A multi-head attention mechanism that projects the input into queries, keys, and values.
|
||
\item A feed-forward network that processes the attention output.
|
||
\item Residual connections and normalization applied after both components.
|
||
\end{itemize}
|
||
|
||
\paragraph{Decoder:}
|
||
|
||
In a sequence-to-sequence Transformer, the decoder generates the target sequence autoregressively, one token at a time, conditioned on the input sequence processed by the encoder.
|
||
|
||
At inference time, decoding begins with a special \emph{start-of-sequence} token. Each token is embedded and augmented with positional encodings to incorporate information about token order, just as in the encoder.
|
||
|
||
Each decoder block consists of three main sublayers:
|
||
|
||
\begin{enumerate}
|
||
\item \textbf{Masked Multi-Head Self-Attention}:
|
||
The decoder first applies multi-head self-attention over its own output sequence.
|
||
To prevent positions from attending to future tokens (which would leak information during training),
|
||
a causal \emph{look-ahead mask} is applied to the attention weights.
|
||
Formally, for input $Y$ to the decoder:
|
||
|
||
\begin{equation}
|
||
\text{MaskedAttention}(Q, K, V) = \text{softmax}\left( \frac{QK^\top}{\sqrt{d_k}} + M \right)V
|
||
\end{equation}
|
||
|
||
where $M$ is a mask matrix with $-\infty$ in positions that should be masked (future positions), and $0$ elsewhere.
|
||
|
||
\item \textbf{Cross-Attention (Encoder-Decoder Attention)}:
|
||
The output from the masked self-attention is used to attend to the encoder’s representations.
|
||
This allows the decoder to incorporate context from the entire source sequence.
|
||
The queries come from the decoder hidden states, while the keys and values come from the encoder outputs:
|
||
|
||
\begin{equation}
|
||
\text{CrossAttention}(Q, K, V) = \text{Attention}(YW^Q, XW^K, XW^V)
|
||
\end{equation}
|
||
|
||
where $X$ is the encoder output, and $Y$ is the decoder’s self-attention output.
|
||
|
||
\item \textbf{Feed-Forward Network and Normalization}:
|
||
The result of the cross-attention layer is passed through a feed-forward network (FFN),
|
||
followed by residual connections and layer normalization, mirroring the encoder architecture:
|
||
|
||
\begin{align}
|
||
y_1 &= \text{LayerNorm}(Y + \text{MaskedMultiHead}(Y)) \\
|
||
y_2 &= \text{LayerNorm}(y_1 + \text{CrossAttention}(y_1, X)) \\
|
||
y_3 &= \text{LayerNorm}(y_2 + \text{FFN}(y_2))
|
||
\end{align}
|
||
\end{enumerate}
|
||
|
||
Finally, the output of the last decoder block is projected into the vocabulary space using a learned weight matrix $W_{\text{vocab}}$
|
||
and passed through a softmax function to produce a distribution over the next possible token:
|
||
|
||
\begin{equation}
|
||
P(\text{next token} \mid \text{previous tokens}) = \text{softmax}(y_3 W_{\text{vocab}}^\top + b)
|
||
\end{equation}
|
||
|
||
During inference, the model generates tokens one at a time, feeding each predicted token back into the decoder to generate the next.
|
||
This continues until a special \emph{end-of-sequence} token is produced.
|
||
|
||
Multiple decoder blocks can be stacked to increase model depth and expressivity.
|
||
While originally developed for machine translation, the Transformer decoder has since been applied to a wide range of tasks,
|
||
including time-series forecasting and biomedical modeling~\cite{wu_deep_nodate,zeng_are_2022}.
|
||
Its ability to model complex, long-range dependencies without recurrence makes it well-suited to domains like biomedical time-series,
|
||
where signals are often irregular and span diverse temporal resolutions.
|
||
|
||
This makes Transformers well-suited for learning long-range temporal dependencies in physiological data,
|
||
such as ovulatory trends spanning multiple days or cycles.
|
||
|
||
\subsubsection{Convolutional Layers as Temporal Feature Extractors}
|
||
For high-resolution time-series data, the input dimensionality can become large,
|
||
especially in models like Transformers that process the entire sequence in parallel.
|
||
This can lead to increased memory consumption and slower training.
|
||
To mitigate this and retain as much information as possible, convolutional layers can be used
|
||
to reduce the sequence length while preserving important local patterns.
|
||
|
||
In this context, one-dimensional convolutions act as learnable filters that slide over the input sequence to extract temporal features.
|
||
Each filter is parameterized to respond to specific local structures in the data, such as peaks, slopes, or short-term motifs.
|
||
By adjusting the \emph{stride}—the step size of the convolution—the model can control the degree of downsampling,
|
||
effectively reducing the number of time steps passed to subsequent layers.
|
||
|
||
Additional dimensionality reduction can be achieved using pooling operations, such as \emph{max pooling}, which retains only the maximum value within a given window.
|
||
These techniques reduce the computational load while maintaining salient information for downstream processing.
|
||
|
||
Figure~\ref{fig:background_convolution_example} illustrates a simple one-dimensional convolution applied to a sequence using a filter of size 3.
|
||
The stride determines how far the filter moves at each step, affecting both the resolution and length of the resulting feature map.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.6\textwidth]{resources/figures/background/background_convolution_example}
|
||
\caption{Example of a simple 1-D convolution on an input sequence.}
|
||
\label{fig:background_convolution_example}
|
||
\end{figure}
|
||
|
||
In summary, time-series modeling offers a range of approaches, each with specific trade-offs.
|
||
RNNs and LSTMs provide explicit sequential modeling but suffer from training inefficiencies.
|
||
Transformers excel at long-range context capture but demand more memory and parallelization.
|
||
Convolutional layers offer efficient local feature extraction and often serve as useful pre-processing stages for both model families. |