Multivariate Financial Time Series Analysis: VAR, cointegration, and state space models
Vector self-regression model
The global integration of the economy and the development of information dissemination have linked financial markets, and price changes in one market can spread quickly to another. Investors holding multiple assets also wanted to know about the relationship between returns on multiple assets. These issues fall within the context of a multi-temporal sequence analysis. We start with this chapter by looking at multiple time series analysis, rather than treating them as individual analyses.
Multi-time series basic concept
Weak Smooth Row
When a multiple time series $r_{t}= {r_{1t}...r_{nt}}$ Call it a plural smooth column if the following conditions are met
$$\begin{cases}E\bardsymbol{r} t=\bardsymbol{mu}\text{not }t\text{){mathrm{= = \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ = = = = = = { { { { { { { { { { { { { { { { { { { { { { {Cov}(\boldsymbol{r}t,\boldsymbol{r}== sync, corrected by elderman == @elder man&You're not gonna get away with this?
It can be seen that the concept of multiple and weak stratification is naturally transformed from the one dollar wide and smooth.Random Process Basis "Soft Process" section
Interrelate Matrix
A single-dollar time series would require only a study of the factors associated with the deviation and lag, but the diversity would need to be considered more.
Let's remember. $$\rho_{ij}(0)=\mathrm{corr}(r_{it},r_{jt})=\frac{\mathrm{Cov}(r_{it},r_{jt})}{\sqrt{\mathrm{Var}(r_{it})\mathrm{Var}(r_{jt})}}=\frac{\Gamma_{ij}(0)}{\sqrt{\Gamma_{ii}(0)\Gamma_{jj}(0)}}$$ To delay the 0-sync multi-time series interconnective matrix, he was a symmetric matrix of all 1 symmetrical elements of the diagonal sequence, studying the relevance of the various sub-series of the multi-time series, which he had modified from the normal synoptic matrix.
To study the lag relationship, we define it.$k$ Weak and smooth sequences $r_t$ Delay$l$ The matrix of the mutual agreement is $$00-$-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-00-0-0-00-0-00-00-0-t-\boldsymbol{\mu})(\boldsymbol{r}$ ^t-l}bardsymbol(m) He's also a natural extension of the one dollar self-conciliation difference, relying on delay rather than time.
Delay in obtaining amendments from them$l$ The interconnective matrix is $$\rho_{ij}(l)=\mathrm{corr}(r_{it},r_{j,t-l})=\frac{\Gamma_{ij}(l)}{\sqrt{\Gamma_{ii}(0)\Gamma_{jj}(0)}}$$ He's not normally symmetrical, and when the lag-related matrix is not zero, we usually call it a pioneer.
The method of calculating the matrix of sample interconnectivity is easy to imagine. $a ..l=\frac1T\sum{t=l+1}^T(\boldsymbol{r}t-\bar{\boldsymbol{r}})(\boldsymbol{r}{tl}-bar(baroldsymbol{r}) ^T$ Sample interconnective matrix can be calculated from the inter-coordinate matrix.
Classification of linear dependencies between time series
The interconnectivity of multiple time series reflects the linear dependencies of the time series, and here is a simple summary chapter.
We've recorded the interconnectivity of multiple time series as $p_l$ The elements are: $r_{ij}(l)$ Then we can give it to you.
- Diagonal elements $r_{ii}(l)$ It's a one-dollar time series. $r_{it}$ ACF
- $p_{ij}(0)$ It's two points. $r_{it},r_{jt}$ Synchronise Linear Relationships
- $p_{ij}(l)$ It's reeling. $r_{it}$ Yeah.$r_{jt}$ And zero is not linear.
Based on the difference. $p_{ij}(l)$ And in the case of a multi-temporal sequence, we can divide it into one.
- $p_{ij}(l)=p_{ji}(l)=0$ For Any$l$ The two sequences are irrelevant.
- $p_{ij}(l)=p_{ji}(l)=0$ For Any $l>It's a zero-dollar split.
- One is not a zero, called a one-way guide and lag.
- Neither of them is zero, called mutual guidance and lag.
Multiple-compositing Test
The one-dollar Ljung-Box white noise test was extended to a variety of situations. Test zero for a multi-series. $$H_0:\boldsymbol{\rho}_1=\cdots=\boldsymbol{\rho}_m=\boldsymbol{0}$$ The opposing assumption is not all zero matrix.
Use test statistics $$Q_k(m)=T^2\sum_{l=1}^m\frac1{T-l}\mathrm{tr}(\hat{\Gamma}_l^T\hat{\Gamma}_0^{-1}\hat{\Gamma}_l\hat{\Gamma}_0^{-1})$$ You can achieve a white noise test similar to the Ljug-Box, and determine if the sequences are white noise.
VAR Model Foundation
VAR Model Structure
The most common of multiple asset-rate joint models is the Vector Autoregression, VAR model, which we give you.$k$Won's$VAR(1)$ The model structure is... $$\bardsymbol{r}0+\boldsymbol{\Phi}\boldsymbol{r}\bardsymbol}t$ of which$\phi_0$ Yes.$k$Other Organiser $\Phi$ Yes.$k$ Array $a_t$ It's a error column. It's usually assumed to be zero.$k$Normal distribution
Consider $k=2$ The model structure is becoming $$\left{\begin{array}{l}r_{1t}=\phi_{10}+\phi_{11}r_{1,t-1}+\phi_{12}r_{2,t-1}+a_{1t}\r_{2t}=\phi_{20}+\phi_{21}r_{1,t-1}+\phi_{22}r_{2,t-1}+a_{2t}\end{array}\right.$$ If $\phi_{12}=\phi_{21}=0$ If you're separated, you can call the two sequences separate.$a_t$ It's not relevant, and we call it non-conformity. Conversely, if the coefficient for determining the separation is not zero, it is called that the two sequences are feedback-to-respond.
Statistics have their own way of explaining the relationship between these feedbacks, when two sequences of feedbacks are made.$a_{1t}$ and $a_{2t}$ It's not relevant, it's called a transmission function. We can adjust it.$r_1$ To adjust.$r_2$ In econometrics, it's called Granger's Causation.
Granger has a more detailed explanation for this: considering a binary sequence ahead of schedule.$l$Step prediction problems, using VAR models and one-dimensional models, respectively, to predict if$r_{2t}$ The two-dollar projection is more accurate than his one-dollar projection.$r_{1t}$ It's the Granger cause. Of course, it could be for the Grande cause.
We do not explain in detail the reasoning behind the prediction error, which is essentially the simplest MSE, and returns to the example before it.
When?$\phi_{12}=0$♪ Time, predict ♪$r_2$ It's needed.$r_1$ ♪ Information, so ♪$r_1$ Yes.$r_2$ The reasons for the Granger are the same as for the other. When the new interest item of the sequence$a_t$ The alignment matrix is not the time of the diagonal, and the two sequences are synchronized, i.e., the transient grange causality.
The way it is, it's all about the way ahead.$VAR(1)$ Research does not allow for the discovery of grange causality in a practical application from such a simple coefficient relationship. But that's enough to understand the model itself.
Simplified structure of VAR
In the model structure used earlier $\Phi$ Reflects dynamic dependencies, and synchronized dependencies are used$a_t$ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . $\Sigma$ . This form is usually called the simplified form of the VAR model, because it does not clearly show the synchronous linear dependency between the spectrospecies.
We can use matrix variations to express the synchronous dependencies in a visible way.$a_t$ . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . $\Sigma$ Exists in Cholesky Disaggregated by $\Sigma = LGL^T$ of which$G$Yes.$k$ The Quarters,$L$It's a lower triangle matrix with an angle of 1. $$\bardsymbol{b}t=\boldsymbol{L}^{-1}\boldsymbol{a}t=(b{1t},\ldots,b{kt})^T$$ 则有 $$\begin{aligned}E\boldsymbol{b}t=&0\\mathrm{Var}(\boldsymbol{b}t)=&\boldsymbol{L}^{-1}\mathrm{Var}(\boldsymbol{a}t)L^{-T}=\boldsymbol{G}\end{aligned}$$ 因此我们可以对原本的VAR模型进行同时左乘$L^{-1}$ 得到 $$\begin{aligned}\boldsymbol{L}^{-1}\boldsymbol{r}{t}=&\boldsymbol{L}^{-1}\boldsymbol{\phi}{0}+\boldsymbol{L}^{-1}\boldsymbol{\Phi}\boldsymbol{r}{t-1}+\boldsymbol{L}^{-1}\boldsymbol{a}{t}\=&\boldsymbol{\phi}{0}^{}+\boldsymbol{\Phi}^{}\boldsymbol{r}{t-1}+\boldsymbol{b}{t}\end{aligned}$$ 他的最后一个子方程为 $$* * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * * $ $ $ * * * * * * * $ $ $ $ $ $ $ $ Because$b_{kt}$ Must be with$b_{ki}$ So this equation is a direct reflection of the simultaneous dependency we call the structure equation.
Simplified forms are commonly used in time series analysis, because
- Simplified forms are easier to estimate;
- At the time of the projection, the synchronized form is not available;
Steady conditions and rectangularness
We analyzed the problem with the stability of the AR model in the online time series analysis. Linear Time Series Analysis The section on "Stable and random time series ARMA / Self-Regression Process (AR)" and uses the feature multi-form thinking to study its stability.
There are similar problems in VAR models, but they are still complex and are not presented here.
VAR(p) model
Here we expand the VAR(1) to VAR(p) model, we call it.$k$The time series obeys.$VAR(p)$ When? $$\bardsymbol{r}t=\bardsymbol{\pi}0+\bardsymbol{Phi}1\boldsymbol{r}{t-1}+\cdots+\boldsymbol{\Phi}p\boldsymbol{r}\bardsymbol{a} The rules for all of these coefficients are not changed.
Yes.$VAR(p)$ , coefficient of model$\Phi$ It's a precursor to each weight, but it's complicated.
VAR Modeling
Estimates and rankings
VAR model modelling also follows largely the repeated trial processes of ranking, model estimation and model testing. One dollar of PACF can be extended to multiple situations to support rankings.
For a real data, we consider the next step progressive VAR model. $$00\ &\boldsymbol{r}_t= \boldsymbol{\phi}_0+\boldsymbol{\Phi}1\boldsymbol{r}{t-1}+\boldsymbol{a}_t \ &\boldsymbol{r}_t= \boldsymbol{\phi}_0+\boldsymbol{\Phi}1\boldsymbol{r}{t-1}+\boldsymbol{\Phi}2\boldsymbol{r}{t-2}+\boldsymbol{a}_t \ &\text{:} \ &\boldsymbol{r}_t= \boldsymbol{\phi}_0+\boldsymbol{\Phi}1\boldsymbol{r}{t-1}+\cdots+\boldsymbol{\Phi}p\boldsymbol{r}{t-p}+wordsymbol{a} t I'm sorry, I'm sorry. Model parameters can be estimated using OLS (minimal 2 times) for each equation, i.e. multi-linear regression problems.
We're the first of them.$i$The equation is estimated to be the difference. $ }{t}^{(i)}=\boldsymbol{r}{t}-\hat{\boldsymbol{\Phi}}{1}^{(i)}\boldsymbol{r}{t-1}-\cdots-\hat{\boldsymbol{\Phi}}{i}^{(i)}\boldsymbol{r}{t-i}$$ 他的协方差矩阵为 $$\hat{\boldsymbol{\Sigma}}i=\frac1{T-(k+1)i-1}\sum^That(boldsymbol{a)}t$ So we can go one by one.$l$Performance of hypothetical tests $H 0:\bardsymboll=\mathbf{0}\leftrightarrow H_a:\boldsymbol{\Phi}Other Organiser Test statistics at $M(1)=- (T-k-\frac{5} \ln\frac{hat{\bardsymbol{\Sigma}{1}|}{|\hat{\boldsymbol{\Sigma}}He was following the C.O. when the hypothesis was set
Or we could use information guidelines like AIC to determine that he needs to use the synoptic matrix of the much-approached margin in the form of $ \tilde(boldsymbol}{i}=\frac{1}{T}\sum{t=i+1}^{T}\hat{\boldsymbol{a}}{t}^{(i)}[\hat{\boldsymbol{a}}You're not gonna get it. The form of the definition of the volume of information is not presented here, and the selection results of these guidelines are not influenced by the matrix.
Model testing
Model differences can be calculated and multiple white noise tests (multiple mixing tests) are performed for the residuals. The multiple-compositing tests of the disability are reduced by using the estimated parameters.$k^2p$, this is the coefficient matrix $\Phi _j, j= 1, 2, \ldots , p$.
If some parameters in the coefficient matrix are fixed to zero, the freedom to be deducted should be calculated in the number of parameters without binding.
Simplified Model
When VAR is measured$k$When larger, the model has many parameters, and the number of parameters in the coefficient matrix is$k^2p$A few. If there is no a priori knowledge requirement that the parameters are not zero, the less significant parameters can be bound to zero and then estimated.
This is consistent with our one-dollar time-series model in practical applications, based on$t=test$ And visualize to fix certain coefficients to zero.
Grange's karma test.
If the model can be simplified to include a coefficient equal to zero for some of the GL/R, then the GL/R test can be performed accordingly. In the binary VAR(1) model, if bound$\phi_{12}(1)=0$The post-module model does not differ significantly from the unbounded model, and$r_{2t}$No, it's not.$r_{1t}$The Granger cause.$p$Levels and$k$The dollar is similar.
To compare unbound and unbound models, the logarithmic test is used to test the statistically obtained amounts closer to the calculator distribution under the zero assumption that the binding parameters are equal to zero.
The test of Grange's causality is based on VAR, and the limitation is that the weights must be smooth, not support the co-ordinated model. So the co-ordination model can't use the function that's here.
Projections
If VAR$p$) Models are known, smooth conditions are met, set${a_t}$It's a separate, stable time series. Use it.$F_t$Other Organiser$t$It's been so long.$r_s,s\leq t$, and then $E (\bardsymbol{a}t|F{t-1})=0$。基于$t$时刻的信息进行超前$I$ 1 trot forecast, projected $$\bardsymbol{r}t(l)=E(\boldsymbol{r}{t+l}|F_t)$$ 当$l=1$时 $$\boldsymbol{r}_t(1)=\boldsymbol{\phi}_0+\boldsymbol{\Phi}_1\boldsymbol{r}_t+\cdots+\boldsymbol{\Phi}p\boldsymbol{r}{t+1-p}$$ 当$l=2$时 $$\begin{aligned}\boldsymbol{r}t(2)=&E(\boldsymbol{r}{t+2}|F_t)\=&\boldsymbol{\phi}_0+\boldsymbol{\Phi}1E(\boldsymbol{r}{t+1}|F_t)+\boldsymbol{\Phi}2\boldsymbol{r}t+\cdots+\boldsymbol{\Phi}p\boldsymbol{r}{t+2-p}\end{aligned}$$ 若记 $$\left.\boldsymbol{r}t(l)=\left{\begin{array}{ll}E(\boldsymbol{r}{t+l}|F_t),&l>0\\boldsymbol{r}{t+l},&l\leq0\end{array}\right.\right.$$ 则超前$l$步预报可以写成 $$r_t(l)=E(r{t+l}|F_t)=\boldsymbol{\phi}0+\sumI'm sorry, I'm sorry. The visible advance multistep projection can be calculated incrementally.
For VAR that meets the condition of stability$(p)$Models, that can be proved. $$\lim_{l\to\infty}\boldsymbol{r}_t(l)=\boldsymbol{\mu}=E\boldsymbol{r}_t$$That's what predicts the average regression.
The prediction error is easy to write. $$\bardsymbol{e}t(l)=\boldsymbol{r}{t+l}-\boldsymbol{r}t(l)=\boldsymbol{r}{t+l}-E(\boldsymbol{r}_{t+l}|F_t)$$
Accomplishment and vector error correction model
Fake return.
Linear regression analysis is one of the most common models of statistics, but if the self-variant and the variable that returns is a time series, then the time series is the one that is the only one that can be used to make a statistical model. The return does not satisfy the basic assumption of the regression analysis: the model error item is distributed independently.
When such a false return problem arises, the return may not be compatible or the standard error estimates and hypothetical tests that coincide with the return result are incorrect. We're here.Linear Time Series Analysis The section on “Regression models with time series errors” describes a more common situation and treatment of false regressions.
We will continue to discuss the design of false returns more closely and to refine the relevant theory.
Linear Time Series Analysis The original sequence is given sufficient differentials to ensure its smoothness in the section on regression models containing time series errors so that the final error series must be in the form of time series, the latter part of this chapter Coordination and analysis It'll be a little more complicated.
Coordination and analysis
Concept of coordination and analysis
For the binary time seriest=(x{1t},x_{2t})^T$,如果$x_{1t}$和$x_{2t}$都是一元单位根过程,但存在非零线性组合$\beta=(\beta_1,\beta_2)$使得 $z_t=\beta_1x_{1t}+\beta_2x_{2t}$弱平稳,则称两个分量$x_{1t}$和$x_{2t}$存在协整关系(cointegration) , $(\beta_1,\beta_2)^T$称为$x t$ integer.
Multiple multiple time series of points can similarly define the concretization relationship, and multiple concretization vectors can be present in multiple cases.
The two-staged Engle and Granger method
Want to look at multiple time series$r_t$The unit root test is required to confirm that both parts are unit root processes, and that there is no unit root after the difference, which is called "single-size"
Second, I'll...$x_{1t}$The blogger says:$x_{2t}$As a variable, as a linear regression, it gets disabled.$e_t$Sequences, and regression factor$\beta_1$, the equation is $$x_{1t}=\beta_0+\beta_1x_{2t}+e_t$$
According to the study by Engle and Granger, the parameters are estimated to be the same at the time the conciliation relationship is formed, but the coefficient estimates are not normal, so the estimate is obtained using a linear minimum of two times the estimate.Point estimates are available, but the t and F tests in the result are not valid。
To verify that the association is established, only a unit root test of the returning defect is required, and when he does not have a unit root, we call the two points a combination. But because...$e_t$ It's a return disability, so we need to use the Phillips-Ouliaris co-program.
The second stage of the two-stage Engle and Granger approach is the need to find all the condensed vectors in a number of situations, which requires the modification of the vector error model (VECM), which we will present later.
VARMA Model
Following the one dollar ARMA model, the VAR model can be extended to VARMA model in the form of $P(B)\bardsymbol{r}t=Q(B)\boldsymbol{a}t$$ 其中 $$\begin{aligned}P(z)=&\boldsymbol{I}-\boldsymbol{\Phi}{1}z-\cdots-\boldsymbol{\Phi}{p}z^{p}\Q(z)=&\boldsymbol{I}+\boldsymbol{\Theta}{1}z+\cdots+\boldsymbol{\Theta}You're not gonna get a job. VARMA has a problem with the same model that can be expressed as different parameter forms, so use VAR as much as possible to avoid VARMA.
Error correction model and reconciliation
Error fixer model
Because in the system, the number of units of non-stable weights is more than the number of units of root (the linear combination can reduce the amount of unit roots not even), so if the difference is calculated for each unit of non-stable weights, it is smooth, but it causes excessive differences.
This excess is a country differential that we will not see in a one-dollar model.
To correct this excess, we propose a vector error correction model.
For a VARMA model if it contains$m$A complication factor.$m$ The following form of error correction (VECM) $$$$$\Delta\boldsymbol}t=\boldsymbol{\alpha}\boldsymbol{\beta}^T\boldsymbol{x}{t-1}+\sum_{j=1}^{p-1}\boldsymbol{\Phi}j^*\Delta\boldsymbol{x}{t-j}+\boldsymbol{a}t+\sum{j=1}^q\boldsymbol{\Theta}j\boldsymbol{a}$ of which$\boldsymbol{\alpha}$and$\beta$Both.$k\times m$The ma is full of matrixes, no roots in the MA.$m$D-time series.$\bardsymbol{y}t=\boldsymbol{\beta}^T\boldsymbol{x}t$是平稳列(没有单位根),$\boldsymbol{\beta}$的每一列都是$\boldsymbol{x}t$的一个协整系数。$\boldsymbol{\Phi}j^$和$\boldsymbol{\alpha},\boldsymbol{\beta}$都依赖于原来的AR部分的系数矩阵 $\boldsymbol{\Phi}j., in relation to: $ \begin{aligned}\bardsymbol{\{j}^{}=&-\sum{i=j+1}^{p}\boldsymbol{\Phi}{i},j=1,2,\ldots,p-1\\boldsymbol{\alpha}\boldsymbol{\beta}^{T}=&\boldsymbol{\Phi}{p}+\cdots+\boldsymbol{\Phi}Other Organiser coefficient$\alpha,\beta$ It's not the only one.
Related uses
The VEPM model is determined using the maximum semblance estimate
The VECM model needs to be tested using Johansen's co-processing test, which is the essence of the test.$\mathbf{\Pi}=\boldsymbol{\alpha}\boldsymbol{\beta}^T$ Yes.$rank(\Pi)$ The test is the number of coordinated relationships.
The estimated VEPM model can be used for prediction. First, you can get it from the model.$\Delta\boldsymbol{x}_t$The prediction of the sequence, then from$\Delta\boldsymbol{x}_t$- I can solve it.$\boldsymbol{x}_t$ . The difference between VEPM and VAR projections is that VEPM permits unit roots and unit roots, and VAR forecasts do not allow unit roots.
VECM is the only time series model we've been learning to allow the root of the unit to exist, which is his unique place.
State spatial model
A brief introduction.
State-spatial models are powerful, flexible and diverse models in the area of time series analysis, which, in conjunction with Kalman filtering techniques, can cover ARIMA models, many non-stable models with external variables, and many models with a different range of variables. More than before (includes)Linear Time Series AnalysisLinear time series models described in ) are more flexible。
R ExtensionstatespacerAnd many models based on linear Goss State spatial models have been achieved, and can be customised. The state space model is a relatively independent knowledge, but it's quite powerful, compared toFinancial time series analysis (one dollar) and the "Version self-repeat model" section Financial time series analysis (one dollar) I'm sure the "Assessment and Vector Errors Model" section is more used than anything else.
As an introduction, we'll start with a local horizontal model. The model is simple, so it can be used to demonstrate the expression and estimate of the state space model. And then we're looking at the whole state space model.
Local horizontal models
Set${y_t,t=1,2,\ldots,T}$For time series, meet the following models
$ \begin{aligned}y t}=&\mu_t+e_t,:{e_t}\sim\mathrm{iidN}(0,\sigma_e^2),:t=1,2,\ldots,n,\\mu_{t+1}=&\mu_t+\eta_t,:{\eta_t}\sim\mathrm{iid♪ I'm not gonna let you go ♪
of which${e_t}$and${\eta_t}$Independent, initial$\mu_1$For a given value or random variable subject to normal distribution, with ${e t,\eta t,t>0}$相互独立。称${\mu_t}$为${y_t}$的水平,模型中${y_t}$可观测而$It's not visible.
This equation is a special case of a linear Gospel State spatial model. We can see from this model structure a similar place to the various time series structures studied earlier.
Of which${\mu_t}$Called State equation ${y_t}$ Called Observation equation ${e_t}$It's observational error. It's instantaneous error or noise.
This model is called a local horizontal model, and it is a special case of a structural time series model.
We can notice. $$y_t-y_{t-1}=\eta_{t-1}+e_t-e_{t-1},$$ The difference between a first-order differential is the sum of random errors in the average value and a first-order lag, which means the original. $y_t$ Obey. $ARIMA(0,1,1)$
Local horizontal models can handle multiple time series, but they are only processed in a single-dollar sequence, and there's nothing special about them.
Filter, Smooth and Forecasting
We continue to use local-level models as examples of the various analytical and modelling techniques of state-of-the-art spatial models, and filtering, smoothing and forecasting are the issues most frequently considered in the state-of-the-spatial models. The following is the text:
- Filter: From${y_1,\ldots,y_t}$ estimate $\mu_t$
- Smooth: From${y_1,\ldots,y_n}$ estimate ${\mu_1,...\mu_n}$
- Forecast: From${y_1,\ldots,y_t}$ estimate $\mu_{t+h}$ So we can give it to you.
- Filtering as$E(\mu_t|y_{1:t});$
- Smoothly desolate as$E(\mu_t|y_{1:n});$
- Forecast to$E(\mu_{t+h}|y_{1:t})$or$E(y_{t+h}|y_{1:t})$。
For local horizontal models, set$\mu_1\sim\mathbb{N}(a_1,P_1)$and separates from the disturbance sequence. By the nature of the normal distribution, the local horizontal model is the Gaussian process, and the condition distribution is still the Gaussian distribution, so...$\mu_t|y_{1:s}$and$y_t|y_{1:s}$Relying on Goss distribution (multiple-normal distribution), the conditionality is expected to be a minimal-equivalence error estimate and a linear one.
$\mu_t$Yes.$y_1:s$The distribution of the conditions below is determined solely by expectations and differences in conditions. Remember$\mu_t|s=E(\mu_t|y_{1:s})$Remember$\Sigma_t|s=Var(\mu_t|y_{1:s})$。
Remember$y_{t|s}=E(y_t|y_{1:s})$I'm sorry. In particular, remember$a_t=E(\mu_t|y_{1:t-1}),P_t=$Var$(\mu_t|y_{1:t-1})$I'm sorry. By the nature of the distribution of Goss, the conditions are non-random. Remember $$v_t=y_t-E(y_t|y_{1:t-1}),$$ That's right.$y_t$The error in making the best prediction, obviously.$Ev_t=0$You... $$F_t=Ev_t^2=\mathrm{Var}(v_t),$$ The country is a country with a large number of different kinds of countries.$v_t$and$y_{1:t-1}$Independence, so there are. $ \begin{aligned}F t}=&\mathrm{Var}(v_t)=E(v_t^2)=E(v_t^2|y_{1:t-1})\=&E[(y_t-E(y_t|y_{1:t-1}))^2|y_{1:t-1}]\=&\mathrm{Var}(y_t|y_{1:t-1}).\end{aligned}$$
Kalman filter.
Kalman filtering is an infernal algorithm.$t=1,2,\ldots$based on$\mu_t|y_{1:t-1}$Distribution of conditions and newly available observations$y_t$Please.$\mu_t|y_{1:t}$Conditional distribution, which equals$\mu_t|(y_{1:t-1},v_t)$The distribution of conditions requires only expectations and differences in the distribution of conditions in Goss.
The blogger adds:$\mu_t|y_{1:t-1}\sim\mathbb{N}(a_t,P_t),v_t\sim\mathbb{N}(0,F_t)$and$y_{1:t-1}$Independence. Attention.${e_t}$and${\eta_t}$Independent, so...${e_t}$and${\mu_t}$Independent, with the expectation of a filter. $$00\ &\mu_{t|t}=E(\mu_t|y_{1:t}) \ &= E(\mu_t|y_{1:t-1},v_t) \ &= E(\mu_t|y_{1:t-1})+E(\mu_t-E\mu_t|v_t) \ &= a_t+\frac{\mathrm{Cov}(\mu_t-E\mu_t,v_t)}{\mathrm{Var}(v_t)}v_t \ &= a_t+\frac{\mathrm{Cov}(\mu_t,v_t)}{F_t}v_t. \end{aligned}$$ 其中 $$\begin{aligned} &\mathrm{Cov}(\mu_t,v_t) \ &=E(mu tv t)\quad(\text{Note}Ev t=0)\ &= E[\mu_t(y_t-a_t)] \ &= E[\mu_t(\mu_t+e_t-a_t)] \ &= E[\mu_t(\mu_t-a_t)]+E[\mu_te_t] \ &= E[\mu_t(\mu_t-a_t)]+0 \ &= E\left{E\left[(\mu_t-a_t)^2|y_{1:t-1}\right]\right} \ &= E{P_t}=P_t. \end{aligned}$$ 化简有 $$E(\mu_t|y_{1:t-1},v_t)=a_t+\frac{P_t}{F_t}v_t,$$ 记 $$K_{t}=\frac{P_{t}}{F_{t}}=\frac{P_{t}}{P_{t}+\sigma_{e}^{2}},$$ 所以有滤波的条件期望为 $$(mu t 1 ,v t) =a t+K tv t, $ That means we'll put $y_1,...,y_t$ Yeah. $\mu_t$ The best forecast (i.e. filtering) formula is broken down into two parts; the first part is:$y_1,...,y_{t-1}$ Yeah. $\mu_t$ The best forecast. The second part is a new pair.$y_t$The best forecast, the latter factor being the Kalman gain.$K_t$The best forecast is linear.
In practice, the Kalman filtering operation is carried out in a round, each of which is structured as follows, and in the cycle structure of the round, the entire filter sequence is obtained. $$\left{aligned}v t=&y_t-a_t,\F_t=&P_t+\sigma_e^2,\K_t=&P_t/F_t,\a_{t+1}=&\mu_{t+1|t}=a_t+K_tv_t,\P_{t+1}=&\Sigma t+1|t}P t(1-K t)+\sigma \eta^2,\mathrm{t=1,2,\ldots, n.\end{aligned}\right.$ Arguments for initial distribution of algorithms $a_{1}$ and $P_1$ The selection has a significant impact on the entire Kalman filter, and we'll be able to introduce it separately.
One step for errors
When we were ahead of Kalman filtering, one step forecast was made and one step error was studied. $$00\ &v_{1}= y_1-a_1, \ &v_2= y_2-a_2=y_2-a_1-K_1(y_1-a_1), \ &v_3= y_3-a_3=y_3-a_1-K_2(y_2-a_1)-K_1(1-K_2)(y_1-a_1), \end{aligned}$$
We can write it in matrix form, like $$\bardsymbol{bardsymbol{(Y n a mathbf )n)$$ 其中 $$K=\begin{pmatrix}1&0&0&\cdots&0\k{21}&1&0&\cdots&0\k_{31}&k_{32}&1&\cdots&0\\vdots&\vdots&\vdots&\ddots&\vdots\k_{n1}&k_{n2}&k_{n3}&\cdots&I'm sorry, I'm sorry. We'll use it later.
Smoothness of Status and Disturbation
Smooth state
In the filter, we want to predict with the observations we have. $\mu_t|y_{1:t}$ When we get all the observations, ${y_1,...y_n}$ Using all observations to estimate $\mu_t$ That's what I get. $\mu_t|Y_n$ This is called a smoothing problem.
We're going to give the local horizontal model's state smooth calculation method without proving it does.
To get what you want.t=\mu{t|n}$,需要先进行卡尔曼滤波求出$a_t,P_t,v_t,F_t,K_t,L_t$,然后令$r n=0$, calculated by inverse inverse: $ \begin{aligned}r=&\frac{v_t}{F_t}+L_tr_t,\\hat{\mu}{t}=&\mu{t|n}=a_t+P_tr_{t-1},:t=n,n-1,\ldots,2,1.\end{aligned}$$ 同理,可以反向递推计算状态平滑方差有 $$\begin{aligned} N_{t-1}=& \frac1{F_t}+L_t^2N_t, \ V_{t}=& \Sigma_{t|n}=P_t-P_t^2N_{t-1},t=n,n-1,\ldots,2,1. \end{aligned}$$
Disturbing Smoothness
We can also estimate the difference between smoothing and smoothing. $e_{t},\eta_t$ The condition distribution, the problem is called disturbance smooth. He can use it to model, to find a mutation point in the state (a leap or change point for local horizontal models equivalent to a level), to find abnormal values for observed errors.
Remember $a }t=E(e_t|y{1:n}),\quad\hat{\eta}t=E(\eta_t|y{1:n}),\mathrm{~}t=1,2,\ldots,n.$$ 因为 $e_t=y_t-\mu_t$ 所以有 $$e_t\left|y_{1:n}\right.\sim\mathrm{N}(y_t-\mu_{t|n},\Sigma_{t|n})=\mathrm{N}(y_t-\hat{\mu}t,V_t).$$ $$\hat{\eta}t=E(\mu{t+1}|y{1:n})-E(\mu_t|y_{1:n})=\mu_{t+1|n}-\mu_{t|n}=\hat{\mu}_{t+1}-\hat{\mu}_t,$$
We can just give the formula there is. $$\begin{gathered} E(e_t|y_{1:n})= \sigma_e^2\left(F_t^{-1}v_t-K_tr_t\right), \ \mathrm{Var}(e_t|y_{1:n})= \sigma_e^2-\sigma_e^4\left(\frac1{F_t}+K_t^2N_t\right), \end{gathered}$$ For the state equation disturbance, $$00\ =& \sigma_\eta^2r_t, \ \mathrm{Var}(\eta_t|y_{1:n})=& \sigma_\eta^2-\sigma_\eta^4N_t,\mathrm{~}t=n,n-1,\ldots,2,1. \end{aligned}$$
Processing and forecasting of missing values
It is difficult for a general time series model to process missing values that appear within the time horizon. A major advantage of the state spatial model is that it is easier to observe missing values.
In Local Horizontal Models, set${y_t}_{t=\ell+1}^{\ell+h}$Missing. The state space model can solve the problem of missing values in a number of ways, using a method that does not change the time step and model form.
Yeah. $t\in{\ell+1,\ldots,\ell+h}$ We can give you the formula of a local horizontal model. $$\mu_t=\mu_{t-1}+\eta_{t-1}=\cdots=\mu_{\ell+1}+\sum_{j=\ell+1}^{t-1}\eta_j,$$
We can give you filter structure as follows: $$00\ ♪ I'm not sure I'm gonna be able to do this ♪& E(\mu_t|Y_\ell)=a_{\ell+1}, \ \mathrm{Var}(\mu_t|Y_{t-1})=& \mathrm{Var}(\mu_t|Y_\ell)=P_{\ell+1}+(t-\ell-1)\sigma_\eta^2, \end{aligned}$$ 于是有递推式 $$\begin{aligned}a_t=&\mu_{t|t-1}=\mu_{t-1|t-2}=a_{t-1},\P_t=&♪ I'm not gonna let you go ♪ I'm sorry. Which means that the Kalman filter we've been doing is still working, for the missing ones.$y_t$ We should take the corresponding.$v_{t}= 0$ At the same time$K_{t}= 0$ That means no Kalman gain.
In fact, the prediction we're making is basically Kalman filter, and the results are the same as those for future values, which are given for missing direct filters.
Selection of primary value distribution parameters and model parameter estimates
Selecting the parameters of the primary distribution
Kalman filters need to be presumed to know. $\mu_1\sim\operatorname{N}(a_1,P_1)$ Actually, one of them.$a_1,P_1$ It's all unknown.
Use filter formula $$00\ ♪ The world is so full of shit ♪& y_1-a_1,\quad F_1=P_1+\sigma_e^2, \ a_{2}=& a_1+\frac{P_1}{F_1}v_1=a_1+\frac{P_1}{F_1}(y_1-a_1) \ \rightarrow & y_1\quad(P_1\to\infty), \ P_{2}=& P_1\left(1-\frac{P_1}{P_1+\sigma_e^2}\right)+\sigma_\eta^2 \ =& \frac{P_1}{P_1+\sigma_e^2}\sigma_e^2+\sigma_\eta^2 \ \rightarrow & The blogger says that the government is not a party to the law. I'm sorry, I'm sorry. And so...$P_1\to\infty$It's like a moment of thought.$y_1$is a non-random defined value, and$\mu_1\sim\mathbb{N}(y_1,\sigma_e^2)$I'm sorry. This initialization method is called proliferation initiation or diffusion a priori. The a priori proliferation is equivalent to a lack of knowledge of the initial state distribution.
Model parameter estimates
Filtering and smoothing are hypothetical model parameters.$\sigma_e^2$and$\sigma_\eta^2$Known. For the estimation of parameters, the maximum semblance method can be used, and the filter algorithm can be used to calculate the apparent function.
State spatial model
All of our previous presentations are related knowledge of local horizontal models, which are a simple exception to linear Gospel spatial models. This section gives a state spatial model, examples of other models that this model can represent, and gives filters, smoothing, forecasting formulas and parameter estimation methods.
- Take care of your reference. R TSA The “State-Spatial Model” section of this section is very important for modelling in the form of model marks, which are followed by the State-Spatial Model.
Many models can be presented as state-spatial models, but researching this expression is not very meaningful in applications.
Linear Goss State Space Model
The state spatial model has many different expressions, according to the formula (Durbin and Koopman 2012), and the linear Goss model is: ^b0460f $$00begin{gathered} \boldsymbol{t=t\boldsymbol{\alpha}t+\boldsymbol{\varepsilon}t,\boldsymbol{\varepsilon}t\sim\mathrm{N}(0,H_t), \ \boldsymbol{\alpha}{t+1}= T_t\boldsymbol{\alpha}_t+R_t\boldsymbol{\eta}_t,\boldsymbol{\eta}_t\sim\mathrm{N}(0,Q_t), \end{gathered}$$ 其中 $$\alpha_1\sim\mathrm{N}(a_1,P_1).$$
of which$y_t$Yes.$t$The time value of the observation is$p\times1$Vector;$\boldsymbol{\alpha}_t$Yes.$t$The state of the time system is non-observable.$m\times1$Random vector, first equation called observation equation, second equation called state equation.
${\boldsymbol{\varepsilon}_t}$and${\boldsymbol{\eta}_t}$The two are independent and distributed white noise columns.$\boldsymbol{\varepsilon}_t$Yes$p\times1$Random vector,$\boldsymbol{\eta}_t$Yes$r\times1$Random vector,$r\leq m$。
Set Matrixs$Z_t,T_t,R_t,H_t,Q_t$The blogger says:$Z_t$and$T_{t-1}$ Allows dependency on \\bardsymbol{y}1,\ldots,\boldsymbol{y}{t-1}$,初始状态$\boldsymbol{\alpha}_1$服从$N(\boldsymbol{a}_1,P_1)$,设$\boldsymbol{a}_1,P_1$已知,$\boldsymbol{\alpha}_1$与${\boldsymbol{\varepsilon}_t}$和$I'm not gonna be able to get a chance to get a chance to get a chance to get a better chance.
Set when parameters are unknown$\psi$Unknown parameter, matrix$Z_t,T_t,R_t,H_t,Q_t$You can rely on unknown parameters.$\psi$。
in the Model$R_t$The government has been able to control the situation.$r=m$Some of the models of teaching materials are not available.$R_t$This one. Organisation$R_t$The good news is,$R_t$It's often a team formation.$I_m$♪ Some columns make up one ♪$m\times r$Matrix, called Selection Matrix, which allows for an error of zero in the equation for some state mass, and$\boldsymbol{\eta}_t$Square Range$Q_t$It could be full of stuff.$r\times r$Stand by, if not.$R_t$Matrix$Q_t$I'm not sure if I'm gonna be happy. If$R_t$It's normal.$m\times r$The matrix, most of the conclusions on the state spatial model, remain valid.
Promulgated state spatial model
This is a succession to the previous section, which extends the stateal Goss spatial model to the state equation, which is still linear Gossic, and the observation equation is distributed in non-Gas, or the observational equation is not linear in relation to the state variable, and it is also extended to the state equation in non-linear, and is distributed in non-Gs.
The generic non-linear, non-Gross State spatial model is in the form of a generic non-linear, non-Gross-state spatial model. $$00\begin{aligned}\bardsymbol{y}{t}\sim&f{t}(\boldsymbol{\alpha}{t};\boldsymbol{\beta}),\\boldsymbol{\alpha}{t+1}\sim&g t}(\bardsymbol {\alpha};\bardsymbol{\theta},\end{aligned} I'm sorry. Such models typically require random simulation methods such as MCMC, sequenced and important samples for filtering, smoothing and estimation.
MASS package model
MARSSIt's a more common R-state space model software package, and he has some agreement on the model format, which we need to understand to facilitate our use of the software package. MARSS is the abbreviation of multiple self-regression spatial models, which are, in fact, linear Gaussian spatial models.
The basic model formula is: $$00\&\boldsymbol{x}{t}=B\boldsymbol{x}{t-1}+\boldsymbol{u}+\boldsymbol{w}{t},\quad&\boldsymbol{w}{t}\sim\mathrm{N}(0,Q),\&\boldsymbol{y}{t}=Z\boldsymbol{x}{t}+\boldsymbol{a}+\boldsymbol{v}{t},\quad&\boldsymbol{v}{t}\sim\mathrm{N}(0,R),\&\bardsymbol0}sim\mathrm{\ (\bardsymbol{\pi},\Lambda).\end{aligned} I'm sorry. Here andFinancial time series analysis (one dollar) The structure of the “Line Gosse State Space Model” section is essentially the same, with only a change of marking. The special feature is the matrix. $B,Z,u,a$ Time change is allowed.
More complex models can also add a section on the impact of external variables to two equations. The parameterization and estimation methods of the MARSS extension package differ considerably from other state-of-the-art spatial model extension packages.
The regression part of the external variable and the models that can be written in each matrix allow for variations $$00\&\boldsymbol{x}{t}=B{t}\boldsymbol{x}{t-1}+\boldsymbol{u}{t}+C_{t}\boldsymbol{c}{t}+\boldsymbol{w}{t},\quad&\boldsymbol{w}{t}\sim\mathrm{N}(0,Q{t}),\&\boldsymbol{y}{t}=Z{t}\boldsymbol{x}{t}+\boldsymbol{a}{t}+D_{t}\boldsymbol{d}{t}+\boldsymbol{v}{t},\quad&\boldsymbol{v}{t}\sim\mathrm{N}(0,R{t}),\&\bardsymbol0}sim\mathrm{\ (\bardsymbol{\pi},\Lambda).\end{aligned} I'm sorry. of which$c_t$It's in the equation.$p$External variable data, which can be entered into one$p\times T$Matrix;
$C_t$is the corresponding regression load matrix, which can contain unknown quantities, if it is non-temporal, if entered as$m\times p$matrices; if time changes, enter as$m\times p\times T$3-D array, with the last subscript for time$t$。
$\boldsymbol d_t$It's in the observation equation.$q$The data on the external variables,$D_t$is the corresponding load matrix.
This is not just a regression model, but also an internal state variable.$x_t$. A time series model with an external variable (return from the variable).
The Herma Model HMM
The Hemama model is similar to the state spatial model, but it follows a chain of horsees, usually in dispersive form. The model is also used extensively, for example, in biological research, model identification, financial modelling, etc.
- (Zucchini, MacDonald, and Langrock 2016): Hidden Markov Models for Time Series - An Introduction Using R. 2nd ed., 2016, CRC Press.
HMM Basic Introduction
Preparatory knowledge
The observational values variable of the Hema model is subject to a simple, marginal distribution in a combination of separate distributions. Set$\delta_1,\ldots,\delta_m$It is a weighted average factor.$p_j(x)$,$j=1,2,\ldots,m$Yes.$m$Density (or probability mass function), $$p(x)=\sum_{j=1}^m\delta_jp_j(x),$$ then$p(x)$is a density (or probability mass function) with a distribution called a stand-alone hybrid distribution or abbreviated distribution. Set$X_j\sim p_j$,$X\sim p$, then $$E(X)=\sum_{j=1}^m\delta_jE(X_j).$$ and $$E(X^k)=\sum_{j=1}^m\delta_jE(X_j^k).$$
We can use the description of the Marseilles. Random Process Basis The "Markov Chain of Dispersed Time" section
The Consort of the Consorture
Set${C_t}$For the chain of horse,${X_t}$For random processes,$X_t$ Yes.$X_1,\ldots,X_{t-1},C_1,\ldots,C_t$The conditions are equal to$X_t$Yes.$C_t$, which is the${X_t}$Follow the Hema model. In fact, the state space model is also the hide-and-map model, but the state equation in the state space model is not normally a discrete horse-map chain.
♪ If the chain is ♪${C_t}$# The only state space #$m$a value, then the model is called$m$Status HMM. Other names of the Hema process, we see names that naturally come to mind.
Set$p_i(x)$Organisation$X_t$Yes.$C_t=i$under conditions, the probability mass function at discrete distribution and the probability density function at continuous distribution.
The simple nature of the Cascade chain.
For the distribution of a dollar, we can directly state: $$P(X_t=x)=[\boldsymbol{u}(1)]^T\Gamma^{t-1}P(x)\mathbf{1}.$$
The distribution of the two dollars is $$00\ &P(X_t=v,X_{t+k}=w) \ &= \sum_{i=1}^m\sum_{j=1}^mu_i(t)p_i(v)\gamma_{ij}(k)p_j(w) \ &= \boldsymbol{u}(t)^TP(v)\Gamma^kP(w)\mathbf{1}. \end{aligned}$$
We can give you the nature of the rectangular. $$00\ E (X t)& =\sum_{i=1}^mu_i(t)E(X_t|C_t=i) \ &== sync, corrected by elderman == @elder man I'm sorry, I'm sorry.
The Cascade chain is a function.
The observation sequence of the Henmar model is$T$Yeah, yeah.$\boldsymbol{X}^{( t) }= ( X_1, \ldots , X_t) ^T$, $\boldsymbol{x}^{( t) }= ( x_1, \ldots , x_t) ^T。( x_1, \ldots , x_T)$ , i.e.$P(\boldsymbol{X}^{(T)}=\boldsymbol{x}^{(T)})$, needs to be$P(X_1=x_1,\ldots,X_T=x_T,C_1=c_1,\ldots,C_T=c_T)$Every one of them$C_t$Item about$c_t$Peace, yes.$T$And all that is to be reconciled.$2T$the product of the entry, so the apparent function of the surface is calculated to be$O(Tm^T)$,$T$It is not feasible to calculate at a larger time; however, there is a general measure of calculation in practice$O(Tm^2)$The algorithm.
We give the apparent function as a matrix. $$L_T=\boldsymbol{\delta}^TP(x_1)\Gamma P(x_2)\Gamma P(x_3)\cdots\Gamma P(x_T)\mathbf{1}.$$ of which$P(x)=\operatorname{diag}((p_1(x),\ldots,p_m(x))),p_j(x)=P(X_t=x|C_t=j)$Not dependent.$t$value.
Maximum apparent estimation method
It's enough to get a simple theory about what this section is about.
Observation values projections, status estimates
The maximum semblance of the values after the given observations is estimated, and then the missing observations can be estimated, the predicted observations, the estimated marsal chain state, etc. This is based on the calculation of the condition distribution.
We don't want the chain to be smooth. $\delta$ Yes. $t=1$ Status$C_1$ Distribution
Conditional distribution of observations
Remember$\boldsymbol{x}^{(-t)}$In the$\boldsymbol{x}^{(t)}=(x_1,\ldots,x_T)^T$Delete$x_t$.$\boldsymbol{X}^{(-t)}$Meanings similar. Consider$X_t$Yes.$\boldsymbol{x}^{(-t)}$condition distribution, which can be used to fill the missing value.
To calculate $$P(X_t=x|\boldsymbol{X}^{(-t)}=\boldsymbol{x}^{(-t)})=\frac{P(X_t=x,\boldsymbol{X}^{(-t)}=\boldsymbol{x}^{(-t)})}{P(\boldsymbol{X}^{(-t)}=\boldsymbol{x}^{(-t)})},$$
And finally, the structure can be written. $$P(X_t=x|\boldsymbol{X}^{(-t)}=\boldsymbol{x}^{(-t)})=\sum_{j=1}^mw_j(t)p_j(x),\mathrm{~}t=1,2,\ldots,T.$$ of which $$w_j(t)=\frac{d_j(t)}{\sum_{k=1}^md_k(t)}.$$
Projected distribution of observations
$$\text{预测分布指条件概率}P(X_{T+h}=x|\boldsymbol{X}^{(T)}=\boldsymbol{x}^{(T)}),\quad\text{可以看成是}X_{T+1},\ldots,X_{T+h}\text{缺失情况下的计算。}$$ At this point, $$P(X_{T+h}=x|\boldsymbol{X}^{(T)}=\boldsymbol{x}^{(T)})=\frac{P(\boldsymbol{X}^{(T)}=\boldsymbol{x}^{(T)},X_{T+h}=x)}{P(\boldsymbol{X}^{(T)}=\boldsymbol{x}^{(T)})},$$ The final projection distribution can be simplified into the form of a mixed distribution of the observed conditions below.
$$P(X_{T+h}=x|\boldsymbol{X}^{(T)}=\boldsymbol{x}^{(T)})=\sum_{j=1}^m\xi_j(h)p_j(x),$$
Decoding
Decoding is a state of restoration based on observations.$C_t$ * The present document was not edited before being sent to the United Nations translation services. $C_t$
The way to decode it, we'll skip it.
State forecast
It can be shown that the state prediction can be decoded at an equal price.
Model selection and diagnosis
Increase in status$m$It changes the alignment, but it will.$m^{2}$Speed increases the number of parameters and there is a risk of over-composed. Some of the special models may streamline the status transfer matrix or condition distribution so that it relies only on a small number of parameters.
The AIC, BIC guidelines can be used to make different models. For the adequacy of model formulation, a false disability can be calculated and a residual diagnosis can be performed.
Model selection using AIC, BIC
Use of information volume guidelines to select models whose ideas need not continue to be repeated
Modelled diagnosis using false disability
After the model is prepared, the adequacy of the alignment needs to be assessed, and the anomalies identified, which are particularly poor. When modelling a normal linear regression model, a model can be used to diagnose the disease; in more general cases, a "false disability" or a "specific defect" can be defined to make model diagnosis.
Set$X$Subject to a continuous distribution, the distribution function is$F(\cdot)$, then$U=F(X)$Obey U(0,1) distribution. Random Variables$X_t$If the observation value is$x_t$, the distribution function is calculated under the presumed model $$u_t=P(X_t\leq x_t)=F_{X_t}(x_t),$$ , and when the model is correct,$u_t$It should be subject to U(0,1) distribution.$u_t$Near zero or one. Because of the conversion of the different distributions to 0 and 1, the observations from the different distributions are comparable.
Set Data As$x_1,\ldots,x_T$, the model is$X_t\sim F_t$,$x_t$Yes.$X_t$And the observations, because of the different distributions, are these$x_t$It's not comparable. Calculate$u_t=F_t(x_t)$,$u_1,\ldots,u_T$For the sake of eveny, these are comparable. Yes, I can.$u_1,\ldots,u_T$The histogram and the QQQ graph, which are not evenly distributed, indicate that the model was wrongly set if there were significant differences with the balanced distribution performance.
The flatness margin is not easily used to identify the anomaly, and the 0.01 and 0.05 fractions are only 0.04 different, and are already very different in the normal distribution. Because we know the distribution of normal, we define the pseudo-psychological. $$z_t=\Phi^{-1}(u_t)=\Phi^{-1}(F_t(x_t)),$$ It is easier to identify abnormalities. The normal-state falseness should be represented in the standard-normal distribution sample when the model is correct. The value of the normal false margin reflects the$x_t$The degree of deviation from the median (not the average) of its distribution. Hetograms, normal QQ Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q Q
The most important nature of the pseudo-disability is that it is distributed in a similar manner as the standard distribution (or the standard normal distribution), and it cannot be assumed that it is independent of each other, and that the falseness is not independent of each other.
Concordant variables and other dependencies
Impacts such as time trends, seasonal items, can be introduced into the model as non-random competitors. Consideration could also be given to the case of a hidden state model being a second or higher chain of horse. The presumption of conditionality could also be relaxed.
HMM with competitive variables
The probability of observing condition distribution parameters or a marzipan chain transfer can be operated on the basis of a competitor. This would still allow for the maximum semblance of estimates. The value of the competitor is known.
HMM based on a second-grade chain
HMM in a continuum
Status Number$m$Sometimes it's hard to choose objectively, when$m$Too many unknown parameters are available when it is large. So sometimes the Hema model of continuity may be more advantageous. This is very close to the state space model.
The semi-hidden mast model.
The state variable is sometimes not accurate by using a first-order chain. The shift to a high-strate marzipan increases the number of parameters. Another extension is the status process, which is half-map chain.
Set$Y_t$Is the status space as${1,\ldots,m}$The Time Zima chain, the state transfer matrix.$\Omega$Angular elements are equal to zero. This is the way to make a difference in the state series when you are two times adjacent.
Set$d_i$It's a probability distribution on a positive integer set, yeah.$i=1,2,\ldots,m$Yes.$m$This is a distribution called the time-suspension distribution. From$Y_t$and${d_i}$Construct Process${C_t}$See below. Every one.$Y_t$Value represents several consistent states.$C_s$, and the number of constants, when$Y_t=i$Timeline Time Distribution$d_i$。
That's how it got here.${C_t}$It is not usually a marina chain, known as the SMC. If${C_t}$All$d_i$It's all geometrical, but it's still a marina chain.
The status of the Hema model$C_t$Replaced with a half-map chain, known as the hidden half-map model (HSMM). The hidden half-marticulation is much more complex than HMM. It is also difficult to introduce compost variables.
HMM can be used to expand the state space of HMM to be approximately any HSM.
HMM for Vertical Data
Existing$K$Every individual, every individual, on a continuous basis.$T$A point of time, and the observation is...${x_{tk},t=1,\ldots,T,i=1,\ldots,K}$I'm sorry. This is called vertical data and economics is called panel (pane) data. It is noted that there is a correlation between multiple observations from the same individual. The same model is used for setting the time series for each individual, but the parameters can vary.
In some cases, it can be assumed.$K$Each sequence depends on a common potential status sequence.$C_t$, the conditions for independentness between the sequences after the given status sequence may be considered for HMM. One example is considering the return on a large number of equities, which are collectively affected by the same potential market position. This can be seen as a multidimensional HMM.
In some cases, it is not possible to assume that there are common state sequences, for example, multiple measurements at different time points for different patients. If it is assumed that the sequences, and the corresponding state, are independent, the function appears to be the product of the apparent functions of each series. If one assumes that some of the parameters in the observation sequence models are the same, the model can be estimated using data from all observation sequences together, increasing the accuracy of the estimate, and this combined approach can model models that cannot be estimated by individual modelling if the data are short.
Changes between individuals can be distinguished by the covariant value.
- Title: Multivariate Financial Time Series Analysis: VAR, cointegration, and state space models
- Author: Hyacehila
- Created at : 2025-09-27 10:50:29
- Link: https://hyacehila.github.io//blog/2025/09/27/multivariate-financial-time-series-analysis-notes/
- License: This work is licensed under CC BY-NC-SA 4.0.