Spatial Data Analysis
Introduction
Introduction of spatial data analysis
Data with spatial coordinates or relative positions are known as spatial data. The classic statistical methodology requires, in most cases, that samples be independent of each other, large samples and repeated repeatedly. Spatial data, on the other hand, generally do not meet the requirements of independence, while spatial heterogeneity exists and cannot be repeated.
The questions in this article can also be addressedRandom process basis: random process definition, digital characteristics and smooth process、Statistical projections: qualitative projections, quantitative projections and extrapolations of trendsHow the concept of a relatively close read together is developed in different contexts.
Migration of classical statistical methods to spatial data needs to be discussed separately. After many years of research, space data have developed their own theoretical systems.The entire spatial data analysis system is based on spatial self-relevance
Origin and development of spatial data analysis
In 1854, John Show, a spatial analysis of cholera data from London, identified sources of infection and became a common source of both spatial data analysis and epidemiology disciplines.
The spatial data analysis originated from spatial interpolation values of mining drilling data, spatial polygonal data analysis from spatially relevant and retrogressive and metrological geology of socio-economic statistics unit data, and spatial point data analysis from ecological sample analysis.
The field of spatial data analysis is being used extensively by many institutions/areas of social science, and machine learning algorithms are being used extensively in spatial data analysis. As the amount of time and space data is generated, we are also entering the age of time and space combined analysis, that is, analysis of space and time data.
Spatial data type
Spatial data are divided into three categories, and we have different spatial data analysis methods for three different spatial data types.
Spatial continuity data
Spatial continuity data is also known as geostatistical data
The representative example is surface temperature analysis, soil drilling distribution data, which can generate continuous data via spatial plug-in values.
Polygon Data
Polygons, also known as face data (real data) or regional data (regional data) He's a graphic information in space, either a ruled (remotely sensed image) or an irregular administrative planning map, which is not continuous, with a specific division of blocks, each of which corresponds to a specific attribute value.
Point Data
Point data. Focus on spatial location, not property values, e.g. spatial distribution of settlements, spatial distribution of outbreak sites.
General form of spatial data
We can record spatial data in the following form. $${z(s):s\in D}$$ of which$D$It's our research area, a subset of the full-coordinate space, an unlimited but generally 2-dimensional plane or 3-dimensional space information. At every point$s$The real value on it is a random variable, all of it.$Z(s)$It's the generality of our population. We get samples of the total space because of the infinite spatial spectrometry (sample)
Based on this general pattern, we can harmonize the original spatial data type.
- For spatial data continuity,$D$Yes.$R^d$A fixed continuum subset
- For polygonal data,$D$Yes.$R^d$A constant set of specs, which no longer have continuous properties.
- For point data,$D$Yes.$R^d$A random subset of,$Z(s)$ It's degraded, no attribute value.
Space data analysis methods
On spatial data
In general, we measure spatial point data and spatial continuity data using point spacing and semi-variant functions. For polygonal data, however, a connectivity matrix is generally used to achieve this. The two forms of expression are different and the thinking is similar.
Spatial data types can be converted to each other and used to express different issues. Forthringham integrates the Kriging model and SAR methods in the polygonal data in a continuous data analysis into a system. We can also convert the number of cases in the original polygonal data into a disease and convert them into point data, or construct the regional equivalence to a continuous data, using different methods for analysis.
Spatial statistical flow
Spatial data analysis has a thinking on the problem that is close to classical statistical analysis, and we all need to extrapolate from the total sample and then the total.
But there is no so-called I.I.D., and when there is a large spatial heterogeneity and the number of samples is insufficient, the location of the sample will significantly influence the statistical inference, i.e., the close connection between the sampling method and the statistical inference.
Model selection
We'll be back with a model by model. This is a general overview.
- There may be space-related overall spatial distribution objects (using Moran)'s $I$Or semi-variant function test)
- There may be spatial heterogeneity.$q$Test)
- Variable acetal issues (different layers or levels)$q$Different, depending on the expertise or maximum$q$(Crowding)
If the model assumes that the subject is of the same general nature, it is the appropriate model, giving the corresponding relationship
- Independent and distributed, using classic statistical methods for research
- The space is highly relevant, but it's not very different.
- Space is highly differentiated, relevant, no explanation variables are not available or explained, and the Sandwich model is appropriate
- Space segmentation and related are very different, with no clear or unaccessible explanation variables, and are broken down into three scenarios
- There are samples of each layer (strata), MSN or P-MSN models.
- No samples for certain layers, BSHADE, P-BSHADE model
- There's only one sample unit with a supporting variable, the SPA model.
- If the explanation variables are clear and accessible, and space is less differentiated and relevant, then the Beyers Level Model (BHM) or multiple regressions are all possible.
We'll go back to the whole story of the system.
Precision assessment
If real values are available, the test is conducted using real values, and consideration may be given to leaving some sample tests; the sample size is insufficient and the classic cross-certification method or one can be used.
If real values are missing, the selection of variables whose nature is close to the target variable is tested; modern real values can be tested by means of historical or future plug-ins without real values.
Overall, improvements in spatial data analysis are not evident in the methodology for precision assessment.
Space Exploration Data Analysis
GIS Profile
The large number of problems in the real world are related to spatial data, and addressing this type of problem requires access to multidimensional spatial coordinates. Traditional statistical software is not very good at this, but, as representative of open source software, R and Python have raised the issue of space data analysis packages, except that there is clearly no more efficient software for processing.
The Geographic Information System (GIS) is a system for spatial data storage, display, management, query, analysis and decision support. This is characterized by the geo-coded data processed as part of the retrieval and processing of information. GIS has a number of specialized software, which need not be limited to traditional statistical analysis software. And...Most GIS also features commonly used spatial data analysis.。
We don't have a presentation here, but we need to study it. GIS is basically the best solution for space exploration data analysis, and SEDA, and R of course, has offered us its own solution.
The most common special GIS is ArcGIS.
GIS principles
The creation of a GIS involves geographical expression, spatial reference, and spatial data models in three parts, and we have here a brief introduction.
Geographical expression of elements
Common geofactory spatial expression of vectors, grids, grids, Voronoi, etc., we understand when we come to the real GIS solution that the pure theory is not clear.
Space reference system
The more common coordinate systems are the geocentric coordinates system, the ball coordinates system and the most common Cartesian coordinates system. The Cartesian coordinate system is the most common of these.
Among the spatial data analysis problems, we need to establish a partial 3D coordinate system, which is normally directly based on the Cartesian coordinate system. There is no need for too much GIS discussion.
In real-world GIS, we most often establish a system of coordinates of the surface, and often global. We know that the Earth is an elliptical body and that the Cartesian coordinate system wants to create a flat coordinate system. So we usually need to do the flattening with a transverse Mercator. Local maps, of course, generally use local plane projection systems. With a projection system, we can easily measure the map, calculate the length, size, length, and properties.
In real-world research, we generally see the Earth as an elliptical sphere, and in order to harmonize standards we introduce a baseline as a measurement benchmark, with the usual baseline being WGS84 ED50 NAD83 for global positioning, European positioning, North American positioning, and 2,000-coordinate systems for domestic use.
The usual projection methods (two concepts with the projection system) have cylinders, cone projection, and azimuth projection, three types of which are unavoidable variations in their angles and directions, and are used in different fields, as shown below. [Spatial data analysis.png]
Space reference is the synthesis of the preceding narratives, and we need to select the reference surfaces that are projected in the coordinates system, and then get a map of the plane.
Spatial data models
For the storage and use of computers, we need to build spatial data models.
In spatial data models, we need to store spatial location data, time data, attribute feature data (coding data) and normal feature data. They usually use vectors to store them.
We will study spatial data models when we present the spatial data analysis in R, which do not need our consideration in most GIS systems with visual interfaces.
General characteristics of space
Spatial data are unique in relation to general data
- Space self-relevance
- Space heterogeneity
- Modified area unit problem
The properties that are relatively close to space tend to be more similar than those that are more distant. It's called Tobler, the first rule of geography, and it's a self-relevance of space. The world is not evenly distributed. As space is divided, the correlation and regression factors change.
These three characteristics are distinguished from the sampling of I.I.D required in classical statistics, and they also give rise to spatial statistics. We'll give you the overall characteristics of the space below.
Space self-relevance
Definitions and impacts
If the vicinity and surrounding areas are more similar to the centre, this is called space-related. If similar values tend to be adjacent to each other, they are referred to as negative spatial correlations.
Non-independent spatial data can affect statistical methods based on independent and distribution assumptions, and in general we can consider
- Scattered samples, reducing correlation between sample points
- Use of space-connected matrix feature values for regression models
- Add space as a variable to the regression model, i.e. the spatial regression method
Space is not only a disadvantage. It makes it possible to [[Spatial data analysis #space plug-in]; at the same time, space regression models can directly use this spatial dependency to improve predictions. ]
Space-related interpretations need to be combined with the value of the indicator itself and knowledge of the relevant area, usually involving experts in the field
Metric
To study space self-relevance, we need to give a space connectivity matrix. $W$ $$W={w_{ij}}$$ When a polygon $i,j$ Take one when you're next to it or zero.
The most common measure of spatial self-relevance is the Moran's I index.$y$Use the next one.$x$ Replace it with a simple mathematical correction to get the following formula.
$$Moran'{\i1 \cH30D3F4}Si=si=si=si}si}si}si}si}si}si^si^si}si}si}si}si}si^si^si^si^si}si}si}si}si}si}si}si}si}si{si{\\si{si{linelinelinesi{si}si}si}si{si}si^si}si^si}si}si^si^si^si}rsi{x{x{x^right^}}}}}$$$$$$$$ We can calculate it. $I$ Mostly in $[-1,1]$ between positives for positives and negatives for negatives and zeros for non-relevance. $x_i,\bar{x}$ The observations are for a point and for the whole.
Moran's I index has its own hypothetical test method.
Space self-relevance can also be measured by a variable function (semi - varigram) $$\gamma\left(h\right)=\frac{1}{2n\left(h\right)}\sum_{s=1}^{n\left(h\right)}\left[x\left(s\right)-x\left(s+h\right)\right]^{2}$$ of which$n(h)$ ♪ To express the distance ♪ $h$ Points logarithmic $x$An image that represents a point observation, and a variable function is generally measured by a variable curve, representing a variable function at a certain distance
We usually set a threshold.$a$ When a variable function is less than $a$ , which is considered relevant, and which is not, the variable function does not assume the test method
Space stratification heterogeneity
Definitions and impacts
Space heterogeneity refers to a variation in the properties over random fluctuations in space, while stratification is a difference in the layers of the layers of the atmosphere of the atmosphere of the atmosphere of the atmosphere of the atmosphere of the Earth.
Space-species heterogeneity is a regularity of space heterogeneity. Heterogeneity itself is the foundation of geography: uniqueness vis-à-vis other locations exists in almost every location.
The presence of spatial stratification is causing the properties of the space environment to be inaccurately depicted as local features, and there are ways in which we use this to be more than a few.
- Classification or zoning, study of regional characteristics
- Local model construction
Space-species heterogeneity is the regularity of the heterogeneity, and stratification modelling may continue to improve the effects of the model that has been created; at the same time, space-plugging values are also dependent on the study of the heterogeneity of the stroposphere. The presence of layers can also help us with sampling.
Metric
The spatial stratification is expressed in classification or partitioning, and is statistically structured as the principle of the smallest difference in the inner layer and the largest difference in the interstory.
So we define the layer of heterogeneity.$q$ We're counting it. $$q=1-\frac{1}{N\sigma^{2}}\sum_{h=1}^{L}N_{h}\sigma_{h}^{2}$$ The range to be taken is $[0,1]$ When it's close to zero, it's the best.
Space General Characteristics Summary
When the whole is independent and stable, we should use classical statistics, and at this point i.i.d. is all data content.
When we test space self-relevance or spatial heterogeneity, we need to use spatial statistics/spatial data analysis for this.
The spatial second-stage smooth assumption is that the attribute value of each point is a random variable, that the mathematical expectations of each point are equal and that the correlation between the two points randomly variable is related only to the distance between them, and not to the absolute location of the two, and that based on the second-stage smooth assumption, we have a space-based self-relevance approach represented by the Kring space plug.
Space segmentation does not satisfy the second-tier smooth assumption, resulting in a spatially differentiating approach represented by geographic detectors and Sandwich space plugs.
All right, all right.
- Classic statistics are used if spatially related and spatially differentiated differences are not significant
- If only space-related is significant, use Kwiding plugs and space regression
- If only space is significant, use Sandwich plugs and stratifications
- If all are significant, use models such as MSN SPA
Space sampling
Spatial sampling is the method of obtaining statistical extrapolation samples. The classic statistical studies of sampling methods are no longer fully applicable in the face of the self-relevance and heterogeneity of spatial data. There is therefore a need to discuss the sampling in spatial statistics separately
In the big data age, the position of spatial sampling and even of the entire sample survey system has declined because we are better able to make complete samples, so here we are briefly presented.
As for the form of sample statistics, we have selected the statistical data for the estimated overall average.
Simple sample in space
A simple sample of space means a number of units drawn from geo-spatial probability, in the form of basic statistical quantities. $$\overline{y}=\frac{1}{n}\sum_{i=1}^{n}y_{i}$$
The exact extract is a point, a sample, an administrative unit acceptable.
Specimen of space systems
The basic idea of system sampling is to extract units at fixed intervals; space system sampling is to spread them slightly as evenly as far as possible into two dimensions. Medium
The basic statistical form is still the same. $$\overline{y}=\frac{1}{n}\sum_{i=1}^{n}y_{i}$$
Space system sampling is more even than simple space sampling, and space system sampling is more appropriate than simple space sampling when the need for interpolation values using space self-relevance is required
Spatial stratification sampling
In the face of spatial heterogeneity, the Specimetric Specimetric Sample and the Sandwich Sample that we're after are more appropriate than the system sample and simple sampling.
The principle of layering is that the difference is the smallest within the layer, the difference between the layers is the largest, and the point where the attribute value is close is divided into the same layer (stratum)
After the layering has been completed, we generally allocate samples to the layers according to some sort of distribution principle. There are common principles.
- Equal distribution of levels
- Distribution according to the number of units in the layer
- Distribution according to the multiplier ratio of standard discrete differences and unit numbers for a layer (with a focus on sampling for the discrete power)
The formula of the statistics remains unchanged. We can calculate the average of the layers.
Sampling of Sandwich
The current sampling method is more often based on sampling of the reporting modules, which in principle should be at least twice the number of reporting units, each of which is drawn at least once. This was detrimental to the compression of sample volumes and the control of sampling costs, and the space Sandwich sample was presented.
The Sandwich sample is an improved layered sampling method that has a good effect on spatial heterogeneity.
First, we still need to stratification and create multiple layers of knowledge. Samples are then distributed to samples based on the knowledge layer, which calculates the average of the statistical volume and the difference. Finally, we match the knowledge layer to the reporting layer, and get the average and the difference between the reporting layer
So we'll do statistical extrapolation in a larger sample and then apply it to smaller units. Go, go, go!
Space Plugin Value
Spatial interpolation is an important part of spatial statistics and can be extrapolated from known points to extract a large number of unknown locations.
Plugin values are also used in classical statistics, but are not widely applied; space-plugging values are more commonly used, and they are useful in practice by extrapolating a small sample value to a larger range of spatial properties.
The method of spatial insertion is closely linked to the overall characteristics of space, and the strength and weakness of the differentness and self-relevance influence our choice of interpolation methods
Nuclear density estimates
The nuclear density estimate is calculated on the basis of the sample point population of the single variable, and its spatially smoothing estimate is calculated.
Use$s$represent any point in space, use$s_i$ It means that the point is known, so it's possible to calculate.$\lambda(s)$ Based on That's a good idea.{\tau}\left(s\right)=\sumI'm sorry, I'm sorry. %1 nuclear function$k$ A predefined inverted U function, which achieves a given$\tau$ As a defined value for smoothing, it is in fact a defined smooth radius. Off$s$The further away, the less it will affect them.
Too big.$\tau$ It's a big effect on the flat local value.
Nuclear density estimates are conducted using space-related self-relevance
Trends Plugin Value
The idea of trend-faced values is straightforward: using a known function to aggregate the distribution, the basic form of the function is determined in advance, with only the unknown parameters estimated.
Trends in the penetration of trends are heavily dependent on the selection of trends; low trends tend to be less likely to produce results, while high trends are more heavily valued
Use of space self-relevance
Inverse distance weight
The value of the feature value of the point to be inserted is the weight of the feature value of the point around it, and weight is inversely proportional to the two-point distance function
Inverse distance is often dependent on manual selection, which tends to lead to a situation where the point of value to be inserted is significantly higher than the surrounding sample point
Use self-relevance to do so
Kriging Method
The Kringing method uses a linear combination of values of several known sample points within the peri-effect range, as follows: $$z_{0}=\sum_{1}^{n}\lambda_{1}z_{1}$$ The Kringing method relies on a second-order smooth assumption.
- The first steps of each point are unknown but consistent.
- The difference between the two points is related to the distance, not to the absolute position.
The Kriging plug is the best way to get it under the same assumptions as above.$\lambda_i$ Make the projections neutral, the calculations.$\lambda_i$ The algebra equation is $$\begin{cases}\sum_{j}^{n}\lambda_{j}C\left(z_{i},z_{j}\right)+\mu=C\left(z_{i},z_{0}\right)\\sum_{j}^{n}\lambda_{j}=1\end{cases}$$ Or... $ \begin{cases}sum^{j}\lambda\gamma\left({i), z \right)+mu=gamma\left({i}, z zright)\sum\ \lambda j}=end{cases}. $ Functions$\gamma$ He and the co-conforming difference are used to measure the degree of correlation.
CoKriging Method
Estimated value$z$& Other Variables$x$ When it's relevant, these compost variables also contain information on the main variable that we can use to supplement our efforts to achieve the following:$z$The estimates, the CoKriging method. $$$220{0}=\sumI'm not gonna let you go. The specific coefficient solver is not repeated here.
Sandwich Plugin
When space relevance is weak, the interpolation method based on space self-relevance cannot be implemented, and here we present Sandwich interpolation, which is the only interpolation method used in this chapter when space is less relevant and space is highly differentiated.
If the correlation is highly different, there is a specific treatment, but not the scope of the study in this section. ] Internal
The steps in calculating the Sandwich plug-in are as follows:
- Layer the target by the minimum intra-group deviation and the largest inter-group deviation to obtain the knowledge layer
- Calculate the average of knowledge layers to the difference
- The report layer of knowledge and more finer grains is superseded to obtain the values of the various reporting units
Method and [(Spatial data analysis#Sandwich sample] The idea is exactly the same, actually, and they do the same thing, combining sampling and extrapolation, not separate.
Space patterns
Space-formulation studies of space differences that are completely beyond random differences fall within the category of [[Spatial Data Analysis #Space Exploration Data Analysis SEDA]], although we do not present our theoretical ideas in those areas.
Space point pattern
Four main methods of identifying the point pattern are different, with input and output, and it is sufficient to select the method according to the form and demand of the data. The data we're analysing are for the [Spatial Data Analysis # Point Data] described earlier.
Sample Analysis
Sample analysis (quadrant anonysis QA) uses a set of square grids to measure the average points and square differences in the grids, and then studies spatial patterns in random, scattered or conglomerate terms using the average and square differences
We usually use VMRs as a specific indicator. $$VRM=\frac{\sqrt{Var(X)}}{\bar{X}}$$ $$VRM\sim \chi^2(n-1)$$ And when the balance is evenly spread,$VRM=0$ When randomly distributed $VRM=1$ When $VRM>At $1, there is a gathering
Periphery Index
The closest neighbor index method (nearest neighbor indicator NNI) judge distribution patterns by point to the nearest distance. The idea is to compare the average distance from the nearest point of view to the nearest point of view of the actual observation and the nearest point of view of the pattern of the random distribution.
The calculation of the nearest distance is $$r=\frac{1}{n}\sum_{i=1}^{n}\min\left(d_{ij}\mid\forall j\right)$$
The nearest range randomly distributed is $Er = 0.5\sqrt{A/n}$ of which$A$ It's the size of the study area.
Define the nearest index NNI has $$NNI = \frac{r}{Er}$$
NNI uses one to divide the line more than one, which means the sample is scattered. Less than one means sample gathering.
Levels
Search for spatially present congregates according to the distance-group approach
Ripley'sK function
The distribution characteristics of the point elements may vary depending on the observation scale, and the concentration of small scales may be random or evenly distributed on a larger scale. Ripley.'sK function can analyse space distribution patterns at any scale, and is therefore the most common method of analysis for space point patterns
Variable Ripley'sK(d) function for distance$d$Average of time within and ratio of event density within the region $$K\left(d\right)=\frac{\sum_{i=1}^{n}N\left(i,d\right)}{n}/\frac{n}{A}=\frac{A}{n^{2}}\sum_{i=1}^{n}N\left(i,d\right)$$ In this formula$n$ The number of incidents in the area is estimated to be about 50 percent.$N(i,d)$Yes and$i$ Distance is$d$ The number of incidents in the range,$A$The area of the study area is, $\lambda=n/A$ It's the spatial density of events.
We can calculate if the situation is evenly distributed.$K(d) = \pi d^2$ So we can construct the following indicators. $$\Delta\left(d\right)=K\left(d\right)-\pi d^{2}或L\left(d\right)=\sqrt{\frac{K\left(d\right)}{\pi}}-d$$ When the indicator above is greater than 0, the point element is concentrated, and less than 0 reflects the spread
Hotspots.
Hotspot study properties are significantly higher than the sub-areas elsewhere
The results of the various hotspot studies are in practice very different.
Gi
The Getis-Ord Gi statistical method identifies hot and cold points by calculating the Gi value for each element. If the Gi value of one region is significantly higher than that of other regions, the region may be considered a hot spot. Conversely, if Gi values are significantly lower than in other regions, it could be a cold spot.
The Gi value formula is $$G_{i}^{*}\left(d\right)=\frac{\sum_{j=1}^{N}w_{ij}\left(d\right)y_{j}}{\sum_{j=1}^{N}y_{j}}$$
All GIS offers Gi values for hotspot assessment, which is likely to be measured by us.$d$Impact
LISA
LISA (local indicator of special association) specifically known as the Moran region's I index to measure spatial self-relevance in local space, i.e. hotspot issues
whose formula is $$I_{i}=\frac{y_{i}-\overline{y}}{S^{2}}\sum_{j}^{n}w_{ij}\left(y_{j}-\overline{y}\right)$$
Spatial scanning statistics SatScan
Space scanning is a method of detecting concentration within the study area using a series of scanning circles. The ratio is calculated by the actual and expected value of the cases in and out of the circle. Depending on the probability distribution of different cases (this is the concentration method for studying diseases), using different seemingly comparable formulas,
Porcelain Appearance Ratio Calculation Formula is $$LR=\left(\frac{c}{\mu}\right)^{c}\left(\frac{C-c}{C-\mu}\right)^{C-c}=\left(\frac{c}{n\frac{C}{N}}\right)^{c}\left(\frac{C-c}{C-n\frac{C}{N}}\right)^{C-c}$$
Space is different.
Space segmentation means studying space stratification heterogeneity, which they can use.$q$ Statistically, we'll study this spatial pattern properties in [Spatial Data Analysis #geographic probe]
Space return
Space self-relevance influences the results of classical linear regression models, and at this point we need to adopt regression models that take into account space relevance
Generic regression model for grid data
The common form of the space regression equation is given by Anselin. Out $0 = \rho W ~ \ varepsilon=\lambda W} \ \ varepsilon+ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ right } } } } } } } } } } } } } } } } } } } \ \ \ \ \ \ \ \ \ } } \ } \ } } } \ \ \ } } } } } } } } } } } } } } } } } } } } } } >$0.00 of which $X$ It's a matrix of variables from the traditional regression model. $y$ It's a observation vector. $W_1$ Connection between reaction samples $p$ It is the coefficient of the spatial lag variable. $W_2$ The spatial connection between the reactional disability can be set up and $W_1$ Exactly the same.
In all, there are three super-parameters in control of the entire regression equation. $\lambda ,\rho, a$
When they took all 0, this was the classic regression equation, which, on the basis of the generic equation, produced two space regression models -- space lag and space error models.
Space lag model
We build on the universal model, the basic form of the space lag model is $$y=\rho Wy+X\beta+\mu $$
This model takes into account the self-relevance of space connections, and actually we're making superparameters at this point. $\lambda = 0$
Space error model
If space dependence is caused by the self-variant that ignores a space impact, then the space error model can model it, and we can then let it be a model for the space impact. $\rho = 0$ Modelling using self-relevance between the disabilities, the basic form of the model is
$$00begin{gathered}
\y=X\beta+\varepsilon \
\varepsilon \lambda W\varepsilon+mu
I'm sorry, I'm sorry.
Anselin suggested that the two models be selected by conducting LMG-error tests, which are more visible, and then reverting to OLS for modelling.
Geographic weighted regression (GWR)
The idea of geo-gravative re-entry is essentially local re-entry, with a local linear re-entry model modelled, and his regression coefficient.$a$ It's no longer a single unit of globality, but a single unit of space.
The idea of a geographically weighted regression is close to [(machine learning introduction and supervision learning#integrated learning# dynamic sorter selection (DCS)], but there are still differences. ]
GWR solvers use a local weighted minimum two-fold regression, which is a function of the geographic distance from the point to be estimated to the other point of observation. The mathematical model is in the form of $$y_{i}=a_{0}\left(u_{i},v_{i}\right)+\sum_{k}a_{k}\left(u_{i},v_{i}\right)x_{ik}+\varepsilon_{i}$$ of which$u_i,v_i$ It's a space coordinate.
Geographic probe
The spatial regression is the process of the variable.$Y$& and self-variant$X$ Linkage, which in fact is also reflected in the consistency of the spatial distribution of variables and self-variant, requires excavation by geographical detectors.
When linear regression models are significant, geographic detectors are necessarily significant, but not necessarily, and as long as there is a correlation between variables, geo-detectors can detect them.
The geo-detectors are subject to space-specific heterogeneity, the basic idea being: So long as the variables have an impact on the variables, there should be consistency in spatial distribution. Here, space can be geospace, attribute space, time classification, and so on.
The geoprospectator contains four detectors.
- Is there space-species heterogeneity, and what factors cause it?
- Variables $Y$ Is there a significant difference?
- $X$ What's the relative importance of it?
- Factors $X$ Yeah. $Y$ Is it independent or is it interactive in any sense?
Space stratification heterogeneity and factor detection
For the first question, we use$q$ Value Metrics $$q=1-\frac{\sum_{h=1}^{L}N_{h}\sigma_{h}^{2}}{N\sigma^{2}}=1-\frac{SSW}{SST}$$
of which$h$ It's a layer number.
$q$ The greater the value, the greater it is.$Y$The more obvious the spatial divide (if the layers are made using Y) When used in the layer $X$ And when it goes on,$q$ The larger the value, the greater the explanation of the variable from the variable.
Risk zone detection
♪ Want to answer the second question ♪$t$ Statistical testing $t (overline{y}{h-1}-\overline{y}{h-2}}=\frac{\overline{Y}{h=1}-\overline{Y}{h=2}}{\left[\frac{Var\left(\overline{Y}{h=1}\right)}{n{h1}}+\frac{Var\left(\overline{Y}{h=2}\right)}{n- I'm sorry, I'm sorry. of which$h$ It's a layer number.
Ecoprospecting
For the second question, we can compare which of the two self-variant variables is more important by constructing F statistics. $$F=\frac{n_{X1}\left(n_{x2}-1\right)SSW_{X1}}{n_{X2}\left(n_{x1}-1\right)SSW_{X2}}$$
Interactive testing
For the fourth question, the method we use is
- Calculate two factors for each. $Y$ Yes. $q$ Value
- The layer that folds two factors into the same layer is new.$q$ Value
- Compare three.$q$ Values.
of which the judgement form is
| The judgement | Interactive |
|---|---|
| $q_{12} < min {q_1,q_2}$ | Non-linear weakening |
| $min {q_1,q_2}<q_{12} < max {q_1,q_2}$ | Single factor is less linear |
| $max {q_1,q_2}<q_{12}$ | Double Factor Enhancement |
| $q_{12} = q_1+q_2$ | Independence |
| $q_{12} > q_1+q_2$ | Non-linear enhancements |
Space and time analysis methods
Time and space analysis requires a combination of time and space data analysis, which is left for future learning, mainly to include
- EOF and Small Wave Analysis
- The biggest entropy in Bayes.
- Bayesian Level Model
- Geo-expulsion Tree
- Genbank Sequence-Sequence Evolution Analysis
- Title: Spatial Data Analysis
- Author: Hyacehila
- Created at : 2026-03-02 04:13:58
- Link: https://hyacehila.github.io//blog/2026/03/02/spatial-data-analysis/
- License: This work is licensed under CC BY-NC-SA 4.0.