461 lines
29 KiB
TeX
461 lines
29 KiB
TeX
%! Author = alex
|
||
%! Date = 3/7/25
|
||
|
||
|
||
\section{Background}\label{sec:background}
|
||
|
||
\subsection{Physiological Background}\label{subsec:physiological_background}
|
||
|
||
\subsubsection{Menstrual Cycle}\label{subsec:menstrual_cycle}
|
||
The menstrual cycle describes the physiological changes in the female body that prepare it for pregnancy.
|
||
It is divided into two phases: the \textbf{follicular phase} and the \textbf{luteal phase}.
|
||
\\
|
||
During the follicular phase, the ovarian follicles mature, and the endometrium (the inner lining of the uterus) thickens
|
||
in preparation for a potential implantation of a fertilized egg.
|
||
Around day 14 of a typical cycle, ovulation occurs, marking the transition to the luteal phase.
|
||
Ovulation refers to the rupture of the mature ovarian follicle and the release of an egg cell into the fallopian tube.
|
||
Figure~\ref{fig:background_basic_female_reproductive_system} illustrates the female reproductive system,
|
||
including the ovaries and the fallopian tubes.
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.4\textwidth]{background_female_reproductive_organs}
|
||
\caption{The basic female reproductive system~\cite{wikimedia_commons_basic_2019}.}
|
||
\label{fig:background_basic_female_reproductive_system}
|
||
\end{figure}
|
||
|
||
Ovulation is triggered by a surge in \textbf{luteinizing hormone (LH)}
|
||
and \textbf{follicle-stimulating hormone (FSH)}, following a peak in estradiol levels.
|
||
As ovulation occurs, estradiol levels drop, progesterone levels begin to rise and a slight increase in body temperature
|
||
(typically around 0.5°C) can be observed.
|
||
This marks the beginning of the luteal phase.
|
||
|
||
During the luteal phase, the endometrium thickens further, creating an optimal environment for embryo implantation.
|
||
LH and FSH levels decrease, while progesterone remains elevated to support endometrial maintenance.
|
||
If fertilization does not occur, progesterone levels drop, leading to the shedding of the endometrial lining along
|
||
with the unfertilized egg.
|
||
This process, known as menstruation, marks the beginning of a new cycle.
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.6\textwidth]{background_menstrual_cycle_physiology}
|
||
\caption{Physiological changes during the menstrual cycle~\cite{pedroso_menstrual_2022}.}
|
||
\label{fig:background_menstrual_cycle_physiology}
|
||
\end{figure}
|
||
|
||
The menstrual cycle typically lasts around 28 days, with ovulation occurring near the midpoint.
|
||
However, variations, particularly in the follicular phase length, are common and can be influenced by factors such as stress, diet, exercise and age~\cite{silberstein_physiology_2000}.
|
||
Figure~\ref{fig:background_menstrual_cycle_physiology} provides a detailed overview of the hormonal and physiological changes throughout the menstrual cycle.
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_labeled_cycle}
|
||
\caption{A cycles temperature curve with its phases and ovulation}
|
||
\label{fig:background_labeled_cycle}
|
||
\end{figure}
|
||
|
||
Figure~\ref{fig:background_labeled_cycle} shows the temperature curve over the course of a menstrual cycle with the
|
||
menstruation, fertile phase and ovulation marked.
|
||
It starts with a menstruation and ends just before the next menstruation.
|
||
The follicular phase starts at the beginning and goes on until the ovulation.
|
||
The luteal phase begins at the ovulation and continues until the next menstruation.
|
||
|
||
Not every cycle results in ovulation—a phenomenon known as anovulation—which leads to a monophasic temperature pattern.
|
||
Anovulation can have various causes, including hormonal imbalances, stress, or underlying health conditions~\cite{rosenfield_adolescent_2013}.
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_anovulatory_cycle}
|
||
\caption{Example of a cycle without an ovulation and the resulting absence of a temperature rise}
|
||
\label{fig:background_anovulation}
|
||
\end{figure}
|
||
|
||
|
||
Anovulation is reflected in temperature data as either an absence of a clear temperature rise or a rise
|
||
that is insufficient in magnitude or duration to be considered a reliable indicator of ovulation.
|
||
Distinguishing between ovulatory and anovulatory cycles is challenging, as the only definitive confirmation of
|
||
successful ovulation in a clinical sense is a positive pregnancy test.
|
||
Even ultrasound imaging can only confirm that an egg was released from its follicle—not whether it was successfully implanted or fertilized.
|
||
Figure~\ref{fig:background_anovulation} shows a cycle that does not have an ovulation, and thus no resulting temperature rise.
|
||
|
||
|
||
To illustrate the diversity of real-world menstrual cycles, Figures~\ref{fig:background_long_cycle} and~\ref{fig:background_short_cycle}
|
||
show examples of cycles that are significantly longer or shorter than a normative 28-day cycle.
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_long_cycle}
|
||
\caption{Example of a long cycle with a length of 111 days}
|
||
\label{fig:background_long_cycle}
|
||
\end{figure}
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_short_cycle}
|
||
\caption{Example of a short cycle with a length of 22 days}
|
||
\label{fig:background_short_cycle}
|
||
\end{figure}
|
||
|
||
These irregularities appear not only on a per-cycle basis, but also across time within the same individual.
|
||
Figures~\ref{fig:background_irregular_cycles} and~\ref{fig:background_regular_cycles}
|
||
show examples of a woman with an irregular and a regular menstrual cycle pattern, respectively.
|
||
Raw body core temperature readings are shown in light blue, with a red line indicating smoothing by local regression.
|
||
Vertical dotted black lines mark the beginning of each cycle.
|
||
The irregular example highlights how multiple parameters can vary between individuals:
|
||
cycle length, timing of ovulation, temperature shift magnitude between phases, and intra-phase temperature variability.
|
||
This multidimensional variability underscores the need for adaptive, data-driven models capable of learning personalized patterns---
|
||
rather than relying on population-wide assumptions.
|
||
For a regular cycle pattern, sophisticated analysis or predictions are often not necessary, as the last ovulation day
|
||
can reliably be used as the next.
|
||
|
||
\begin{figure}[htb]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_irregular_cycle_example}
|
||
\caption{Example of a woman with irregular menstrual rhythm}
|
||
\label{fig:background_irregular_cycles}
|
||
\end{figure}
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_regular_cycle_example}
|
||
\caption{Example of a woman with regular menstrual rhythm}
|
||
\label{fig:background_regular_cycles}
|
||
\end{figure}
|
||
|
||
\subsubsection{Fertility Prediction}\label{subsubsec:fertility_prediction}
|
||
Throughout the menstrual cycle, the chance of fertilization varies significantly.
|
||
An egg cell released from the ovary during ovulation, can be fertilized for up to 24 hours.
|
||
However, since male sperm cells can survive up to 6 days inside the female reproductive tract,
|
||
the fertile window is typically defined as the five days before ovulation until one day after ovulation~\cite{dunson_day-specific_1999}.
|
||
Research by~\citeauthor{dunson_day-specific_1999} has shown that the highest chance of fertilization is around one day before ovulation,
|
||
as illustrated in Figure~\ref{fig:background_pregnancy_chance}.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.7\textwidth]{background_pregnancy_chance_over_time}
|
||
\caption{Chance of fertilization depending on the day of the menstrual cycle.
|
||
The highest chance is around one day before ovulation~\cite{dunson_day-specific_1999}.}
|
||
\label{fig:background_pregnancy_chance}
|
||
\end{figure}
|
||
|
||
It is important to note that fertility prediction is inherently dependent on ovulation prediction.
|
||
Since the probability of conception is tightly linked to ovulation timing, the accuracy of fertility prediction methods
|
||
is constrained by the precision of ovulation detection.
|
||
This relationship underscores the necessity of developing reliable ovulation prediction models,
|
||
as even small inaccuracies can significantly impact fertility assessments.
|
||
|
||
\subsubsection{Physiological Signs of Ovulation}\label{subsubsec:physiological_signs}
|
||
Several physiological signs correlate with ovulation and can be used for prediction.
|
||
As shown in Figure~\ref{fig:background_menstrual_cycle_physiology}, these include hormonal fluctuations (LH and FSH surges) and
|
||
changes in body temperature.
|
||
Additionally, variations in cervical mucus consistency, salivary ferning patterns,
|
||
and electrical resistance of the skin and vaginal mucosa have been observed~\cite{silberstein_physiology_2000}.
|
||
|
||
Among these, ultrasonography provides the most accurate confirmation of ovulation by detecting follicular changes,
|
||
but it is costly and requires specialized equipment.
|
||
Hormone measurements in urine and blood are widely used and available in at-home test kits,
|
||
but they require frequent testing.
|
||
Temperature-based methods, particularly body temperature tracking, offer a non-invasive alternative by
|
||
detecting the slight temperature rise that follows ovulation.
|
||
Advances in wearable technology have further enabled continuous and automated temperature monitoring,
|
||
improving accessibility and usability~\cite{alexander_fertilitatsmonitoring_2014, luo_detection_2020, yu_tracking_2022}.
|
||
|
||
\subsection{Data Source and Characteristics}\label{subsec:data_background}
|
||
This study is based on a dataset collected from users of the \emph{OvulaRing}~\cite{noauthor_ovularing_nodate},
|
||
an intravaginal wearable sensor developed by VivoSensMedical GmbH, located in Leipzig, Germany~\cite{noauthor_vivosens_nodate}.
|
||
The device continuously records intravaginal core body temperature at 5-minute intervals.
|
||
The sensor itself measures approximately 1\,cm $\times$ 1\,cm $\times$ 2\,cm and is embedded in a silicone ring with a diameter of 5\,cm for ease of use.
|
||
It pairs with a smartphone via Bluetooth to synchronize and upload recorded data to a secure database.
|
||
Figure~\ref{fig:background_ovularing} shows an image of the ring attached to its silicone ring.
|
||
|
||
The product has been on the market for over a decade, resulting in an extensive longitudinal dataset of menstrual cycles.
|
||
Cycle boundaries are defined by self-reported menstruation, which users manually log in the accompanying app to mark the beginning of each cycle.
|
||
|
||
Thanks to a battery life of at least six months, the device supports continuous monitoring of long and irregular cycles, enabling the capture of highly variable menstrual patterns.
|
||
|
||
A known limitation of manual cycle annotations is the potential for misalignment.
|
||
Intermediate bleeding events unrelated to menstruation (e.g., ovulatory spotting or irregular shedding) or missing menstruation entries can lead to ambiguous cycle definitions.
|
||
Therefore, all user-entered cycle starts undergo manual review to reduce annotation errors.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.3\textwidth]{ovularing}
|
||
\caption{The OvulaRing sensor attached to its silicone ring~\cite{noauthor_ringpng_nodate}.}
|
||
\label{fig:background_ovularing}
|
||
\end{figure}
|
||
|
||
In addition to temperature measurements, the database includes contextual metadata such as age, height, weight, and optional user-entered markers.
|
||
These markers provide further physiological context and may include information about intermediate bleeding, sexual intercourse, or positive pregnancy tests.
|
||
|
||
At the time of writing, the dataset contains approximately 65{,}000 annotated cycles, comprising more than 350 million individual temperature measurements.
|
||
|
||
\subsubsection{Dataset Summary}
|
||
For the present study, the dataset was reduced to approximately 40{,}000 cycles after filtering out entries that were incomplete,
|
||
contained hardware-related anomalies, or fell outside a reasonable cycle length range.
|
||
|
||
Very short cycles typically result from incorrect cycle start entries or premature termination of temperature recordings.
|
||
Extremely long cycles are often due to data entry errors or pregnancy-related recordings,
|
||
where the sensor was worn continuously throughout gestation—sometimes producing sequences up to nine months long.
|
||
|
||
While such cases may still contain useful information, they were excluded from this analysis to avoid complications in preprocessing and labeling.
|
||
In most instances, only a small portion of these extended cycles contributes meaningfully to the study objectives.
|
||
The cutoff values for cycle length are 10 and 150 days, respectively.
|
||
|
||
The cleaned dataset has 40{,}266 menstrual cycles from 6{,}245 users.
|
||
The median number of cycles per user is 4 (IQR: 2--8) and the median cycle length is 28 days (IQR: 26--32).
|
||
|
||
The average data density—defined as the fraction of available measurements out of the theoretical maximum of 288 measurements per day—is 0.90.
|
||
This corresponds to an average data availability of 90\% per cycle, with an average loss of 10\%.
|
||
|
||
It should be noted, that both ovulation day and anovulation estimates are based on retrospective algorithmic inference, not human labels.
|
||
Further details are provided in section~\ref{sec:methodology}.
|
||
|
||
\subsubsection{Irregularities and Confounding Factors}
|
||
Core body temperature is influenced by various factors unrelated to the menstrual cycle.
|
||
Illnesses—especially those involving fever—can significantly affect temperature patterns.
|
||
This poses a challenge for any analysis relying on temperature data, as one of the key physiological indicators of ovulation is a post-ovulatory temperature rise (see Section~\ref{subsubsec:physiological_signs}).
|
||
|
||
Figure~\ref{fig:background_fever_cycle} shows an example of a cycle where an illness caused a marked increase in temperature.
|
||
This event is particularly problematic because the fever-induced rise occurs just before the expected ovulatory shift, potentially confounding ovulation detection.
|
||
Distinguishing illness-related changes from cycle-related ones requires models that are sensitive to context and robust to outliers.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_fever_cycle}
|
||
\caption{Example of a cycle affected by illness, showing a fever-induced temperature rise shortly before the expected ovulatory shift.}
|
||
\label{fig:background_fever_cycle}
|
||
\end{figure}
|
||
|
||
Another common irregularity arises when users temporarily remove the sensor during menstruation, typically for hygiene reasons—even though the device is safe to wear continuously.
|
||
This behavior frequently results in missing data at the start of each cycle.
|
||
Depending on the total cycle length, this gap can represent a significant portion of the cycle's data.
|
||
|
||
Figure~\ref{fig:background_menstruation_data_gap} shows an example cycle with a data gap during menstruation.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.9\textwidth]{background_menstruation_data_gap}
|
||
\caption{Example of a cycle with a data gap at the beginning, caused by sensor removal during menstruation.}
|
||
\label{fig:background_menstruation_data_gap}
|
||
\end{figure}
|
||
|
||
%TODO: more stats!
|
||
|
||
\subsubsection{Privacy}
|
||
The dataset used in this study contains sensitive personal health information and is handled with strict privacy safeguards.
|
||
All data is pseudonymized and processed exclusively on encrypted devices, ensuring that no identifiable information can be traced back to individual users.
|
||
VivoSensMedical does not share user data with third parties; the data is used solely for internal research and product improvement efforts that directly benefit users at no additional cost.
|
||
|
||
\subsection{Technical Background}\label{subsec:technological_background}
|
||
|
||
\subsubsection{Time Series Analysis}\label{subsubsec:time_series_analysis}
|
||
Time series analysis is a fundamental tool for studying sequential data that evolves over time.
|
||
Unlike other data types, time series data has an inherent temporal order, where each data point is associated
|
||
with a timestamp, capturing its dependence on past values.
|
||
Time series analysis typically serves two main goals:
|
||
Understanding the underlying mechanisms that lead to the observed data and predicting future data points based on the
|
||
historical information and potentially external factors~\cite{cryer_time_2008}
|
||
\\
|
||
Time series analysis encompasses various methods, ranging from simple statistical models to complex deep learning architectures.
|
||
Classical methods
|
||
|
||
|
||
In the following, we will introduce the two most common approaches used for machine learning on time series data,
|
||
LSTMs and transformers.
|
||
|
||
\subsubsection{RNN and LSTM Networks}\label{subsubsec:lstm_networks}
|
||
Recurrent Neural Networks (RNNs) are extensions of classical neural networks that incorporate cyclic connections between neurons.
|
||
These recurrent connections allow the network to retain information from previous inputs by feeding the hidden state
|
||
from a prior time step into the current one, enabling a form of temporal memory.
|
||
|
||
In practice, this means that input data is processed sequentially, one step at a time.
|
||
At each step \(t\), the input \(x_t\) is combined with the previous hidden state \(h_{t-1}\) to produce a new
|
||
hidden state \(h_t\), which contributes to the output \(o_t\).
|
||
This process allows the network to learn temporal dependencies and model sequential data effectively~\cite{medsker_recurrent_1999}.
|
||
|
||
However, this approach has the downside that the model cannot explicitly control how it remembers or forgets information at each step,
|
||
limiting its ability to manage long-term dependencies.
|
||
During training via backpropagation, the weights of a neural network are updated based on the partial derivatives of the loss function.
|
||
The more propagations required (e.g., in deeper networks), the more multiplications are needed to compute gradients for earlier weights.
|
||
|
||
RNNs are typically trained using \textit{backpropagation through time} (BPTT), in which gradients are propagated through many time steps,
|
||
leading to numerical instability.
|
||
If the gradients shrink exponentially, the model suffers from the \emph{vanishing gradient} problem;
|
||
if they grow exponentially, it results in \emph{exploding gradients}~\cite{hochreiter_vanishing_1998}.
|
||
In both cases, learning is significantly impaired.
|
||
|
||
Exploding gradients can often be mitigated using techniques such as \emph{gradient clipping},
|
||
where the magnitude of the gradient is capped---typically within a range of \([-1, 1]\)---to stabilize training.
|
||
|
||
Figure~\ref{fig:rnn_unfolded} illustrates the unfolded structure of an RNN across three time steps.
|
||
This technique, known as \emph{unfolding}, clarifies how sequential inputs update the hidden state and generate
|
||
outputs at each time step.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.8\textwidth]{recurrent_neural_network_unfold}
|
||
\caption{Schematic diagram of the unfolded structure of a recurrent neural network~\cite{fdeloche_english_2017}}
|
||
\label{fig:rnn_unfolded}
|
||
\end{figure}
|
||
|
||
To address the vanishing gradient problem and enable better long-term memory,
|
||
\emph{Long Short-Term Memory} (LSTM) networks were introduced by \citeauthor{hochreiter_long_1997}~\cite{hochreiter_long_1997}.
|
||
LSTMs extend the RNN architecture by incorporating a memory cell and a series of gates that regulate the flow of
|
||
information: the \emph{forget gate}, the \emph{input gate}, and the \emph{output gate}.
|
||
|
||
\begin{itemize}
|
||
\item The \emph{forget gate} determines which information from the previous cell state should be discarded.
|
||
\item The \emph{input gate} controls what new information is added to the cell state.
|
||
\item The \emph{output gate} selects relevant parts of the current cell state to produce the output and the next hidden state.
|
||
\end{itemize}
|
||
|
||
Each gate employs a sigmoid activation function to regulate the flow of information,
|
||
allowing LSTMs to preserve and update memory over long sequences.
|
||
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.6\textwidth]{lstm_cell_diagram}
|
||
\caption{Diagram of a LSTM cell showing the flow of information \cite{chevalier_english_2018}}
|
||
\label{fig:lstm_architecture}
|
||
\end{figure}
|
||
|
||
Figure~\ref{fig:lstm_architecture} illustrates the architecture of an LSTM cell.
|
||
On the left, the previous cell state \(C_{t-1}\) represents the long-term memory,
|
||
while the hidden state \(h_{t-1}\) encodes the short-term memory from the preceding time step.
|
||
The current input \(x_t\) is processed together with these states to update the cell.
|
||
The resulting new cell state \(C_t\) and hidden state \(h_t\) are passed forward to the next time step or used to produce the model’s output.
|
||
Within the cell, the forget gate determines how much of the previous cell state \(C_{t-1}\) is retained.
|
||
The input gate updates the cell state with new information derived from the current input and previous hidden state.
|
||
Finally, the output gate controls how much of the updated cell state contributes to the hidden state \(h_t\),
|
||
which is passed on to the next time step or used for prediction.
|
||
|
||
LSTMs are widely used in biomedical applications due to their capacity to handle sequences of variable length and complexity.
|
||
In the context of ovulation prediction, where hormonal patterns exhibit periodicity but also irregularity,
|
||
LSTMs are well-suited to learn relevant time-dependent signals from sequential physiological measurements.
|
||
|
||
While powerful, LSTMs can be computationally intensive and sensitive to hyperparameter tuning.
|
||
Therefore, they are often compared with alternative architectures,
|
||
including simpler feedforward networks and more recent attention-based models,
|
||
to evaluate trade-offs in performance, interpretability, and computational cost.
|
||
|
||
The next section introduce the \emph{Transformer} architecture, a more recent alternative that forgoes
|
||
recurrence in favor of attention mechanisms.
|
||
|
||
\subsubsection{Transformer Models}\label{subsubsec:transformer_models}
|
||
|
||
Transformer models are a class of neural architectures that use \emph{self-attention}
|
||
to model dependencies in sequential data without relying on recurrence~\cite{vaswani_attention_2017}.
|
||
Unlike recurrent neural networks (RNNs), Transformers process input sequences in parallel,
|
||
allowing them to model relationships between any pair of input tokens or timesteps directly.
|
||
This mitigates the limitations of recurrent models, such as long-term memory constraints
|
||
and vanishing gradients.
|
||
|
||
Originally introduced for machine translation, Transformers have proven broadly applicable to
|
||
various sequence modeling tasks due to their flexibility, scalability, and strong performance
|
||
on complex temporal patterns.
|
||
|
||
At the core of the Transformer is the attention mechanism, which enables the model to compute
|
||
context-aware representations by weighing the importance of different input positions for each output.
|
||
This is achieved through \emph{scaled dot-product attention}, where queries, keys, and values are
|
||
linearly projected from the input and used to compute attention scores.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.4\textwidth]{background_transformer_architecture}
|
||
\caption{The Transformer - architecture for an encoder-decoder model~\cite{vaswani_attention_2017}}
|
||
\label{fig:background_transformer_architecture}
|
||
\end{figure}
|
||
|
||
|
||
Figure~\ref{fig:background_transformer_architecture} illustrates the original encoder-decoder model introduced by~\citeauthor{vaswani_attention_2017}.
|
||
|
||
The Transformer architecture consists of two components: an \emph{Encoder} and a \emph{Decoder}.
|
||
|
||
\paragraph{Encoder:}
|
||
|
||
The encoder is responsible for encoding the input into a contextualized representation.
|
||
In the case of machine translation, this input would be a sentence in the source language.
|
||
|
||
|
||
The input tokens are first mapped to dense continuous vector representations (embeddings).
|
||
Since the attention mechanism permutation-invariant---that is, it does not inherently encode the order of tokens in the sequence---
|
||
\emph{positional encodings} are added to the token embeddings to provide information about the token positions in the sequence.
|
||
|
||
Without positional encoding, repeated tokens such as `The` would be indistinguishable
|
||
to the model regardless of their location, even if they play different syntactic or semantic roles.
|
||
Positional encodings, often based on sinusoidal functions, inject a unique position-dependent signal
|
||
into each token, enabling the model to distinguish between identical tokens in different positions.
|
||
|
||
Inside each encoder block, \emph{Multi-Head-Attention} is applied to the inputs.
|
||
Multi-Head-Attention extends the regular attention mechanism, by adding multiple attention heads that focus on different parts of the embeddings.
|
||
Each head does the scaled dot-product attention independently on its slice of the data.
|
||
The outputs of all heads are then concatenated and combined via a linear projection.
|
||
|
||
The outputs of the attention mechanism are then processed in a feed forward network to allow for a non-linear projection.
|
||
Residual connections for both the attention and the feed forward allow for better gradient flow and model stability.
|
||
|
||
\paragraph{Decoder:}
|
||
|
||
In a sequence-to-sequence Transformer, the decoder generates the output sequence autoregressively,
|
||
using the contextualized representation produced by the encoder.
|
||
|
||
At inference time, generation begins with a special \emph{start-of-sequence} token.
|
||
Like the encoder, the decoder embeds its inputs and augments them with positional encodings
|
||
to retain information about token order.
|
||
|
||
To ensure that the model does not access future tokens during training,
|
||
a \emph{look-ahead mask} is applied within the decoder’s self-attention mechanism.
|
||
This masking ensures that each position can only attend to earlier positions in the sequence,
|
||
preventing information leakage.
|
||
This component is referred to as \emph{masked multi-head self-attention}.
|
||
|
||
Following the masked self-attention, the decoder incorporates information from the encoder
|
||
via a \emph{cross-attention} layer.
|
||
Here, the decoder’s hidden states act as queries, while the encoder’s outputs serve as keys and values.
|
||
This allows the decoder to condition its predictions on the entire encoded input sequence.
|
||
|
||
The output of the cross-attention layer is passed through a position-wise feed-forward network
|
||
and further normalization and residual connections, analogous to the encoder blocks.
|
||
|
||
For tasks such as machine translation, the final decoder outputs are linearly projected
|
||
to the target vocabulary size, and a softmax function is applied to produce a probability distribution
|
||
over possible next tokens.
|
||
|
||
During inference, tokens are sampled sequentially from this distribution and fed back into the decoder for the next prediction step.
|
||
This process continues until a special \emph{end-of-sequence} token is generated, indicating that the model has completed the output sequence.
|
||
|
||
Depending on the use case and data complexity, multiple encoder and decoder layers can be stacked
|
||
to increase model capacity and improve predictive performance.
|
||
|
||
While originally developed for machine translation, the Transformer architecture has since been applied
|
||
successfully to a range of tasks, including time-series forecasting and biomedical data analysis(\cite{wu_deep_nodate,zeng_are_2022}).
|
||
Its ability to model long-range dependencies without recurrence makes it particularly suited for biomedical time-series,
|
||
where signals may be irregular, noisy, or span varying temporal scales.
|
||
|
||
\subsubsection{Convolutional Layers as Temporal Feature Extractors}
|
||
For high-resolution time-series data, the input dimensionality can become large,
|
||
especially in models like Transformers that process the entire sequence in parallel.
|
||
This can lead to increased memory consumption and slower training.
|
||
To mitigate this and retain as much information as possible, convolutional layers can be used
|
||
to reduce the sequence length while preserving important local patterns.
|
||
|
||
In this context, one-dimensional convolutions act as learnable filters that slide over the input sequence to extract temporal features.
|
||
Each filter is parameterized to respond to specific local structures in the data, such as peaks, slopes, or short-term motifs.
|
||
By adjusting the \emph{stride}—the step size of the convolution—the model can control the degree of downsampling,
|
||
effectively reducing the number of time steps passed to subsequent layers.
|
||
|
||
Additional dimensionality reduction can be achieved using pooling operations, such as \emph{max pooling}, which retains only the maximum value within a given window.
|
||
These techniques reduce the computational load while maintaining salient information for downstream processing.
|
||
|
||
Figure~\ref{fig:background_convolution_example} illustrates a simple one-dimensional convolution applied to a sequence using a filter of size 3.
|
||
The stride determines how far the filter moves at each step, affecting both the resolution and length of the resulting feature map.
|
||
|
||
\begin{figure}[htbp]
|
||
\centering
|
||
\includegraphics[width=0.6\textwidth]{background_convolution_example}
|
||
\caption{Example of a simple 1-D convolution on an input sequence.}
|
||
\label{fig:background_convolution_example}
|
||
\end{figure}
|
||
|