Patent
US 12,288,389 B2Patent
Atlas literature
Patent
US 12,288,389 B2Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1 is a block/flow diagram of an exemplary system for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 2 is a block/flow diagram of a method for generating videos with a new motion or content, in accordance with embodiments of the present invention;
FIG. 3 is a block/flow diagram illustrating applying three regularizers on dynamic and static latent variables to enable the representation disentanglement, in …
FIG. 4 is a block/flow diagram of computing the priors during training, in accordance with embodiments of the present invention;
FIG. 5 is a block/flow diagram of exemplary equations for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 6 is a block/flow diagram of an exemplary practical application for learning disentangled representations of vid- eos, in accordance with embodiments of …
FIG. 7 is a block/flow diagram of exemplary Internet-of- Things (IoT) sensors used to collect data/information for learning disentangled representations of …
FIG. 8 is an exemplary practical application for learning disentangled representations of videos, in accordance with embodiments of the present invention;
FIG. 9 is an exemplary processing system for learning disentangled representations of videos, in accordance with embodiments of the present invention; and
FIG. 10 is a block/flow diagram of an exemplary method for learning disentangled representations of videos, in accor- 5 dance with embodiments of the present …
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
A method for learning disentangled representations of videos, the method comprising: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The method of claim 1, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The method of claim 1, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
A non-transitory computer-readable storage medium comprising a computer-readable program for learning dis-entangled representations of videos, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for 10 a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The non-transitory computer-readable storage medium of claim 6, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
A system for learning disentangled representations of videos, the system comprising: a memory; and one or more processors in communication with the memory configured to: feed each frame of video data into an encoder to produce a sequence of visual features; pass the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; apply Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger represen-tation disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approxi-mate a conditional distribution to facilitate minimi-zation of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenate the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The system of claim 11, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The system of claim 11, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
Patent
Atlas literature
Patent
US 12,288,389 B2Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1 is a block/flow diagram of an exemplary system for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 2 is a block/flow diagram of a method for generating videos with a new motion or content, in accordance with embodiments of the present invention;
FIG. 3 is a block/flow diagram illustrating applying three regularizers on dynamic and static latent variables to enable the representation disentanglement, in …
FIG. 4 is a block/flow diagram of computing the priors during training, in accordance with embodiments of the present invention;
FIG. 5 is a block/flow diagram of exemplary equations for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 6 is a block/flow diagram of an exemplary practical application for learning disentangled representations of vid- eos, in accordance with embodiments of …
FIG. 7 is a block/flow diagram of exemplary Internet-of- Things (IoT) sensors used to collect data/information for learning disentangled representations of …
FIG. 8 is an exemplary practical application for learning disentangled representations of videos, in accordance with embodiments of the present invention;
FIG. 9 is an exemplary processing system for learning disentangled representations of videos, in accordance with embodiments of the present invention; and
FIG. 10 is a block/flow diagram of an exemplary method for learning disentangled representations of videos, in accor- 5 dance with embodiments of the present …
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
A method for learning disentangled representations of videos, the method comprising: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The method of claim 1, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The method of claim 1, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
A non-transitory computer-readable storage medium comprising a computer-readable program for learning dis-entangled representations of videos, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for 10 a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The non-transitory computer-readable storage medium of claim 6, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
A system for learning disentangled representations of videos, the system comprising: a memory; and one or more processors in communication with the memory configured to: feed each frame of video data into an encoder to produce a sequence of visual features; pass the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; apply Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger represen-tation disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approxi-mate a conditional distribution to facilitate minimi-zation of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenate the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The system of claim 11, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The system of claim 11, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
Patent
Atlas literature
Patent
US 12,288,389 B2Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1 is a block/flow diagram of an exemplary system for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 2 is a block/flow diagram of a method for generating videos with a new motion or content, in accordance with embodiments of the present invention;
FIG. 3 is a block/flow diagram illustrating applying three regularizers on dynamic and static latent variables to enable the representation disentanglement, in …
FIG. 4 is a block/flow diagram of computing the priors during training, in accordance with embodiments of the present invention;
FIG. 5 is a block/flow diagram of exemplary equations for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 6 is a block/flow diagram of an exemplary practical application for learning disentangled representations of vid- eos, in accordance with embodiments of …
FIG. 7 is a block/flow diagram of exemplary Internet-of- Things (IoT) sensors used to collect data/information for learning disentangled representations of …
FIG. 8 is an exemplary practical application for learning disentangled representations of videos, in accordance with embodiments of the present invention;
FIG. 9 is an exemplary processing system for learning disentangled representations of videos, in accordance with embodiments of the present invention; and
FIG. 10 is a block/flow diagram of an exemplary method for learning disentangled representations of videos, in accor- 5 dance with embodiments of the present …
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
A method for learning disentangled representations of videos, the method comprising: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The method of claim 1, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The method of claim 1, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
A non-transitory computer-readable storage medium comprising a computer-readable program for learning dis-entangled representations of videos, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for 10 a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The non-transitory computer-readable storage medium of claim 6, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
A system for learning disentangled representations of videos, the system comprising: a memory; and one or more processors in communication with the memory configured to: feed each frame of video data into an encoder to produce a sequence of visual features; pass the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; apply Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger represen-tation disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approxi-mate a conditional distribution to facilitate minimi-zation of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenate the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The system of claim 11, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The system of claim 11, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
Patent
Atlas literature
Patent
US 12,288,389 B2Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1 is a block/flow diagram of an exemplary system for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 2 is a block/flow diagram of a method for generating videos with a new motion or content, in accordance with embodiments of the present invention;
FIG. 3 is a block/flow diagram illustrating applying three regularizers on dynamic and static latent variables to enable the representation disentanglement, in …
FIG. 4 is a block/flow diagram of computing the priors during training, in accordance with embodiments of the present invention;
FIG. 5 is a block/flow diagram of exemplary equations for learning disentangled representations of videos, in accor- dance with embodiments of the present …
FIG. 6 is a block/flow diagram of an exemplary practical application for learning disentangled representations of vid- eos, in accordance with embodiments of …
FIG. 7 is a block/flow diagram of exemplary Internet-of- Things (IoT) sensors used to collect data/information for learning disentangled representations of …
FIG. 8 is an exemplary practical application for learning disentangled representations of videos, in accordance with embodiments of the present invention;
FIG. 9 is an exemplary processing system for learning disentangled representations of videos, in accordance with embodiments of the present invention; and
FIG. 10 is a block/flow diagram of an exemplary method for learning disentangled representations of videos, in accor- 5 dance with embodiments of the present …
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
A method for learning disentangled representations of videos, the method comprising: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The method of claim 1, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The method of claim 1, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
A non-transitory computer-readable storage medium comprising a computer-readable program for learning dis-entangled representations of videos, wherein the computer-readable program when executed on a computer causes the computer to perform the steps of: feeding each frame of video data into an encoder to produce a sequence of visual features; passing the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; applying Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for 10 a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger representa-tion disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approximate a conditional distribution to facilitate minimization of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenating the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The non-transitory computer-readable storage medium of claim 6, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
A system for learning disentangled representations of videos, the system comprising: a memory; and one or more processors in communication with the memory configured to: feed each frame of video data into an encoder to produce a sequence of visual features; pass the sequence of visual features through a deep convolutional network to obtain a posterior of a dynamic latent variable and a posterior of a static latent variable; apply Jensen-Shannon divergence for a penalty of static embedding and maximum mean discrepancy for a penalty of dynamic embeddings, and simultaneously minimizing a Wassertein distance and a Kullback-Liebler divergence regularization to trigger represen-tation disentanglement; sampling static and dynamic representations from the posterior of the static latent variable and the posterior of the dynamic latent variable, respectively, with a general sample-based mutual information upper bound employed with a neural network to approxi-mate a conditional distribution to facilitate minimi-zation of a disentangled representation learning (DRL) objective and to reduce dependency of static and dynamic embeddings; and concatenate the static and dynamic representations to be fed into a decoder to generate reconstructed sequences.
The system of claim 11, wherein mutual information of a static factor and a dynamic factor is introduced as a regulator.
The system of claim 11, wherein a set of orthogonal directions is learned in a latent space of generative adver-sarial networks (GANs) to produce meaningful transforma-tions in an image space and to define a disentangled pro-jected latent space by its transpose.
