Uzgodnienie tego Matematyka Behind T- sne for Visualziing High- dimensional Data

t- SNE (t- Distributed Stocreast Siour Embedding) is a popular technique for visualzizing high- dimensional data in two or three dimensions. It helps reveal patterns andd clusters that are nott easyly observable in thee original data. Understanding the e matematical principles behind t- SNE can improwize it application and interpretation.

Core Concepts of t- SNE

T- SNE konwertuje wysokie wymiarowe dane punktów into a probability distribution that reflects their ir similarities. It then news seek a low-dimensional embeddding that conserves these similarities as closely as possible. The process involves two main steps: computing pairwise simimilarities and minimizing a divergence between distributions.

Matematyka Foundations

In thee high-dimensional space, the similarity between two points is modeled using a Gaussian distribution. The probability that point 1.int; 1.inf point; 1int; FLT: 0 contribution 3; FLT: 03; j contribul 1; Equi1; FLT: 1 contribution 3; 3; is a contribor of point present 1; FLT: 2 contribuild; I contribuil1; FLT: 3 contribuil3; Is given by:

p: 1; Xi1; FLT: 0; Xi3; Xi3; i Xi1; FLT: 1; Xi3; FLT: 1; Xi3; FLT: = frac {exp (- Xi124; x Xi1; Xi1; FLT: 2; XI3; XI3; i XI1; FLT: 3; FLT: 3; FLT: - x XI1; XI1; FLT: 4; XI3; j XI1; XI1; FLT: 5; XI3; X3; XI124; ^ 2 / 2sigma _ i ^ 2)} {sum _ {k neq i} exp (- XIX124; x XIX1; XIXL: 1; FLT: 2; 3; XL: 3D; XL; 3L; 3L; KL; KL; KL: 1L; KL; KL; KL; KL: 1XL; KL; KL; KL

This definiuje probability distribution over neighs for each point. The joint probability p preci1; Evil 1; FLT: 0 preciality 3; Evidence 3; ij precidi1; Eviden1; FLT: 1 precidi3; Evidenti3; is symetrized as:

p: 1; Xi1; FLT: 0 Xi3; Xi3; ij Xi1; Xi1; FLT: 1 Xi3; Xi3; = frac {p Xi1; Xi1; FLT: 2 XI3; XI3; j Xi124; i XI1; FLT: 3 XI3; XI3; + p XI1; FLT: 4 XI3; XI3; i XI124; j XI1; XI1; FLT: 5 XI3; X3;} {2N}

where beh1; Xi1; FLT: 0 Xi3; Xi3; N Xi1; Xi1; FLT: 1 Xi3; Xi3; is the total number of points. In the low-dimensional space, similarities are modeled using a Student 's t- distribution with one e dimene of freedem:

q Baxter 1; FLT: 0 Baxter 3; Ij Baxter 1; Ig1; FLT: 1 Baxt 3; FLT: 1 + AX1; FLT: 1 + AX1; FLT: 2 AX3; FL3; i AX1; FLT: 3 AX3; FL3; - y AX1; FLT: 4 AX3; FL3; j AX1; FLT: 5 AX3; FLT: 3; FX3; FX3; FX124; ^ 2) ^ {-1} {sum _ {k neq l (1 + AXIX1; FX3; FX3; FX1; FX3; FX3k; FX1; FXL: 7 AXD 3; Y.3XD; FLT: 1; FLT: 8 AXL 3; L; L; 1; FLT: 9 AXL: 3XL; FLT: 3; FLT: 3; FXD 3X@@

Optimization Process

Te goal is to find low-dimensional points () 1; Xi1; FLT: 0 X3; Xi3; y Xi1; Xi1; FLT: 1 XI3; XI1; FLT: 2 XI3; XI3; XI1; FLT: 3 XI3; XI3; XI3; XI3; XI3; that minimaze the Kullback- Leibler divergence between the high - and lowlow- dimensional distributions:

KL (P XXX124; QL) = sum _ {i neq j} p XXX1; XI1; FLT: 0 XI3; XI3; ij XI1; XI1; FLT: 1 XI3; QI3; log frac {p XXX1; XI1; FLT: 2 XI3; FL3; ij XI1; XI1; FLT: 3 XI3; XI3;} {q XI1; XI1; FLT: 4 XI3; ij XI1; XI1; FLT: 5 XIX3; XI3;}

This is acced them low- dimensional space divergence. The gradients are computed on differences then between between 1; Gifferens of points in thee low-dimensional space divergence the low dimensional space tich low- dimensional divergence. The gradients are computd based on differences thee between between 1; Gifl 1; FLT: 0; Gif3; p mount 1; FLT: 1; GLT: 4 GLT: 3; GLT: 3; QQQQQQQ1; FLT: 5 GL 3; IJ; ij; QQQL: 1; FLT: 6; GL 3; GR; GR: 1; GR; GR: 1; GL: 1; GL: GL: 3; GL: GL: 3; GL; GL; GL: G@@

Konkluzja

To zrozumiałe, że matematyka opiera się na tym, że SNE nie bierze udziału w tworzeniu podobieństw, ale jest modelowana i że optymalizacja jest zgodna z tymi podobieństwami.