Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is an illustration of an algorithm correcting per- spective distortion of a selfie image; 30
FIG. 2
FIG. 2 is a flow diagram of a pre-trained 3D face GAN executing the algorithm and having a pipeline;
FIG. 3
FIG. 3B is a flow diagram of a second part of the pipeline 35 that improves the face latent code and camera parameters of the 3D face GAN;
FIG. 4
FIG. 4 is a flow chart illustrating a method of correcting perspective distortion of a selfie image and generating an improved selfie image using the pipeline:
FIG. 5
FIG. 5 is an illustration including a comparison of the 45 selfie image, ground truth, an image generated by a previous method, and the processed selfie image …
FIG. 6
FIG. 6 includes illustrations of the processed selfie images generated by the 3D face GAN as a function of respective 50 selfie images for different camera …
FIG. 7
FIG. 7 is an illustration including a comparison of the 3D face GAN utilizing measurements determined from a face, or camera’s focal length, or both as …
FIG. 8
FIG. 8 is an illustration including a comparison of the 3D face GAN utilizing a reference image as additional input during improving the face latent code and …
FIG. 9
FIG. 9 is an illustration including a comparison of the 3D face GAN utilizing a burst of images taken when the user moves the camera around the user’s face;
FIG. 10
FIG. 10 is an illustration including a comparison of the 3D face GAN utilizing a depth image as additional input during 65 improving the face latent code and …
FIG. 11
FIG. 11 is a block diagram of electronic components of a mobile device configured for use with the pipeline and method of
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
5 independent · 14 dependent
1
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance.
2
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein processing the input image comprises cropping and aligning the face.
3
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
5
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein fine tuning the gen-erator further comprises: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and using gradient descent to improve parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
7
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein the improved image has a longer camera-to-face distance than the input image.
8
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known iris size of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
9
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known pupillary distance of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
13
Independentpre-trained 3D face generative adversarial network (GAN)
A pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, configured to: process an input image including a face; improve face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tune the generator; and generate an improved image of the face with reduced face distortion due to setting a longer camera-to-face dis-tance.
14
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein the pro-cessing of the input image comprises cropping and aligning the face.
15
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
17
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein fine tuning the generator is calculated by: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and B₂ using gradient descent to optimize parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
19
Independentpre-trained 3D face generative adversarial network (GAN)
A non-transitory computer readable storage medium that stores instructions that when executed by a processor cause the processor to process an image using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer by per-forming the steps of: processing the image including a face; improving face latent code and camera parameters of the 3D face GAN by; freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; tuning a generator of the 3D face GAN; and generating an improved image of the face with reduced face distortion due to a longer camera-to-face distance. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
pre-trained 3D face generative adversarial network (GAN)
discriminatordiscriminator
neuralrendererneural renderer
generatorgenerator
Reported properties
Performance values and ranges asserted in the specification or claims.
Property
Value
Material
Thickness
20–60 cm
—
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2009/0059041 A12009/0059041 A1 3/2009 Kwon
US 2022/0138455 A12022/0138455 A1 5/2022 Nagano et al.
US 2023/0245330 A12023/0245330 A1 * 8/2023 Nadir........................ G06T 5/60examiner
US 2023/0252714 A12023/0252714 A1 * 8/2023 Bradley................ G06T 15/205examiner
US 2024/0193891 A12024/0193891 A1 6/2024 Markhasin et al.
Patent
Atlas literature
Patent
US 12,614,260 B2
SELFIE PERSPECTIVE UNDISTORTION BY 3D FACE GAN INVERSION
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is an illustration of an algorithm correcting per- spective distortion of a selfie image; 30
FIG. 2
FIG. 2 is a flow diagram of a pre-trained 3D face GAN executing the algorithm and having a pipeline;
FIG. 3
FIG. 3B is a flow diagram of a second part of the pipeline 35 that improves the face latent code and camera parameters of the 3D face GAN;
FIG. 4
FIG. 4 is a flow chart illustrating a method of correcting perspective distortion of a selfie image and generating an improved selfie image using the pipeline:
FIG. 5
FIG. 5 is an illustration including a comparison of the 45 selfie image, ground truth, an image generated by a previous method, and the processed selfie image …
FIG. 6
FIG. 6 includes illustrations of the processed selfie images generated by the 3D face GAN as a function of respective 50 selfie images for different camera …
FIG. 7
FIG. 7 is an illustration including a comparison of the 3D face GAN utilizing measurements determined from a face, or camera’s focal length, or both as …
FIG. 8
FIG. 8 is an illustration including a comparison of the 3D face GAN utilizing a reference image as additional input during improving the face latent code and …
FIG. 9
FIG. 9 is an illustration including a comparison of the 3D face GAN utilizing a burst of images taken when the user moves the camera around the user’s face;
FIG. 10
FIG. 10 is an illustration including a comparison of the 3D face GAN utilizing a depth image as additional input during 65 improving the face latent code and …
FIG. 11
FIG. 11 is a block diagram of electronic components of a mobile device configured for use with the pipeline and method of
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
5 independent · 14 dependent
1
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance.
2
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein processing the input image comprises cropping and aligning the face.
3
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
5
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein fine tuning the gen-erator further comprises: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and using gradient descent to improve parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
7
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein the improved image has a longer camera-to-face distance than the input image.
8
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known iris size of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
9
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known pupillary distance of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
13
Independentpre-trained 3D face generative adversarial network (GAN)
A pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, configured to: process an input image including a face; improve face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tune the generator; and generate an improved image of the face with reduced face distortion due to setting a longer camera-to-face dis-tance.
14
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein the pro-cessing of the input image comprises cropping and aligning the face.
15
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
17
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein fine tuning the generator is calculated by: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and B₂ using gradient descent to optimize parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
19
Independentpre-trained 3D face generative adversarial network (GAN)
A non-transitory computer readable storage medium that stores instructions that when executed by a processor cause the processor to process an image using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer by per-forming the steps of: processing the image including a face; improving face latent code and camera parameters of the 3D face GAN by; freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; tuning a generator of the 3D face GAN; and generating an improved image of the face with reduced face distortion due to a longer camera-to-face distance. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
pre-trained 3D face generative adversarial network (GAN)
discriminatordiscriminator
neuralrendererneural renderer
generatorgenerator
Reported properties
Performance values and ranges asserted in the specification or claims.
Property
Value
Material
Thickness
20–60 cm
—
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2009/0059041 A12009/0059041 A1 3/2009 Kwon
US 2022/0138455 A12022/0138455 A1 5/2022 Nagano et al.
US 2023/0245330 A12023/0245330 A1 * 8/2023 Nadir........................ G06T 5/60examiner
US 2023/0252714 A12023/0252714 A1 * 8/2023 Bradley................ G06T 15/205examiner
US 2024/0193891 A12024/0193891 A1 6/2024 Markhasin et al.
Patent
Atlas literature
Patent
US 12,614,260 B2
SELFIE PERSPECTIVE UNDISTORTION BY 3D FACE GAN INVERSION
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is an illustration of an algorithm correcting per- spective distortion of a selfie image; 30
FIG. 2
FIG. 2 is a flow diagram of a pre-trained 3D face GAN executing the algorithm and having a pipeline;
FIG. 3
FIG. 3B is a flow diagram of a second part of the pipeline 35 that improves the face latent code and camera parameters of the 3D face GAN;
FIG. 4
FIG. 4 is a flow chart illustrating a method of correcting perspective distortion of a selfie image and generating an improved selfie image using the pipeline:
FIG. 5
FIG. 5 is an illustration including a comparison of the 45 selfie image, ground truth, an image generated by a previous method, and the processed selfie image …
FIG. 6
FIG. 6 includes illustrations of the processed selfie images generated by the 3D face GAN as a function of respective 50 selfie images for different camera …
FIG. 7
FIG. 7 is an illustration including a comparison of the 3D face GAN utilizing measurements determined from a face, or camera’s focal length, or both as …
FIG. 8
FIG. 8 is an illustration including a comparison of the 3D face GAN utilizing a reference image as additional input during improving the face latent code and …
FIG. 9
FIG. 9 is an illustration including a comparison of the 3D face GAN utilizing a burst of images taken when the user moves the camera around the user’s face;
FIG. 10
FIG. 10 is an illustration including a comparison of the 3D face GAN utilizing a depth image as additional input during 65 improving the face latent code and …
FIG. 11
FIG. 11 is a block diagram of electronic components of a mobile device configured for use with the pipeline and method of
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
5 independent · 14 dependent
1
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance.
2
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein processing the input image comprises cropping and aligning the face.
3
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
5
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein fine tuning the gen-erator further comprises: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and using gradient descent to improve parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
7
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein the improved image has a longer camera-to-face distance than the input image.
8
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known iris size of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
9
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known pupillary distance of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
13
Independentpre-trained 3D face generative adversarial network (GAN)
A pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, configured to: process an input image including a face; improve face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tune the generator; and generate an improved image of the face with reduced face distortion due to setting a longer camera-to-face dis-tance.
14
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein the pro-cessing of the input image comprises cropping and aligning the face.
15
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
17
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein fine tuning the generator is calculated by: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and B₂ using gradient descent to optimize parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
19
Independentpre-trained 3D face generative adversarial network (GAN)
A non-transitory computer readable storage medium that stores instructions that when executed by a processor cause the processor to process an image using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer by per-forming the steps of: processing the image including a face; improving face latent code and camera parameters of the 3D face GAN by; freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; tuning a generator of the 3D face GAN; and generating an improved image of the face with reduced face distortion due to a longer camera-to-face distance. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
pre-trained 3D face generative adversarial network (GAN)
discriminatordiscriminator
neuralrendererneural renderer
generatorgenerator
Reported properties
Performance values and ranges asserted in the specification or claims.
Property
Value
Material
Thickness
20–60 cm
—
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2009/0059041 A12009/0059041 A1 3/2009 Kwon
US 2022/0138455 A12022/0138455 A1 5/2022 Nagano et al.
US 2023/0245330 A12023/0245330 A1 * 8/2023 Nadir........................ G06T 5/60examiner
US 2023/0252714 A12023/0252714 A1 * 8/2023 Bradley................ G06T 15/205examiner
US 2024/0193891 A12024/0193891 A1 6/2024 Markhasin et al.
Patent
Atlas literature
Patent
US 12,614,260 B2
SELFIE PERSPECTIVE UNDISTORTION BY 3D FACE GAN INVERSION
Patent drawings and their descriptions. Click a drawing to enlarge it.
FIG. 1
FIG. 1 is an illustration of an algorithm correcting per- spective distortion of a selfie image; 30
FIG. 2
FIG. 2 is a flow diagram of a pre-trained 3D face GAN executing the algorithm and having a pipeline;
FIG. 3
FIG. 3B is a flow diagram of a second part of the pipeline 35 that improves the face latent code and camera parameters of the 3D face GAN;
FIG. 4
FIG. 4 is a flow chart illustrating a method of correcting perspective distortion of a selfie image and generating an improved selfie image using the pipeline:
FIG. 5
FIG. 5 is an illustration including a comparison of the 45 selfie image, ground truth, an image generated by a previous method, and the processed selfie image …
FIG. 6
FIG. 6 includes illustrations of the processed selfie images generated by the 3D face GAN as a function of respective 50 selfie images for different camera …
FIG. 7
FIG. 7 is an illustration including a comparison of the 3D face GAN utilizing measurements determined from a face, or camera’s focal length, or both as …
FIG. 8
FIG. 8 is an illustration including a comparison of the 3D face GAN utilizing a reference image as additional input during improving the face latent code and …
FIG. 9
FIG. 9 is an illustration including a comparison of the 3D face GAN utilizing a burst of images taken when the user moves the camera around the user’s face;
FIG. 10
FIG. 10 is an illustration including a comparison of the 3D face GAN utilizing a depth image as additional input during 65 improving the face latent code and …
FIG. 11
FIG. 11 is a block diagram of electronic components of a mobile device configured for use with the pipeline and method of
Claims
Claims define the patent's legal scope. Independent claims stand alone; dependent claims (nested) narrow them. Click a claim to expand its dependents.
5 independent · 14 dependent
1
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance.
2
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein processing the input image comprises cropping and aligning the face.
3
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
5
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein fine tuning the gen-erator further comprises: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and using gradient descent to improve parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
7
Dependent← claim 1pre-trained 3D face generative adversarial network (GAN)
The method of claim 1, wherein the improved image has a longer camera-to-face distance than the input image.
8
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known iris size of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
9
Independentpre-trained 3D face generative adversarial network (GAN)
A method of image processing using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, comprising the steps of: processing an input image including a face; improving face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer; obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge; fine tuning the generator; and generating an improved image of the face with reduced face distortion by setting a longer camera-to-face dis-tance, wherein a face-to-camera distance is estimated by a known pupillary distance of the face, wherein the estimated face-to-camera distance is used to improve the face latent code and camera parameters.
13
Independentpre-trained 3D face generative adversarial network (GAN)
A pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer, configured to: process an input image including a face; improve face latent code and camera parameters of the 3D face GAN by: freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; fine tune the generator; and generate an improved image of the face with reduced face distortion due to setting a longer camera-to-face dis-tance.
14
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein the pro-cessing of the input image comprises cropping and aligning the face.
15
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein improving the face latent code and camera parameters comprises: obtaining a depth map and a normal map, and rendering photorealistic face images based on a random face latent and initialization of the camera parameters; calculating losses; and using gradient descent to improve the face latent code and camera parameters until the losses converge.
17
Dependent← claim 13pre-trained 3D face generative adversarial network (GAN)
The pre-trained three-dimension (3D) face generative adversarial network (GAN) of claim 13, wherein fine tuning the generator is calculated by: using the pre-trained 3D face GAN to obtain a depth map and a normal map and render photorealistic face images; calculating losses; and B₂ using gradient descent to optimize parameters of the generator and the neural renderer of the 3D GAN until the losses converge.
19
Independentpre-trained 3D face generative adversarial network (GAN)
A non-transitory computer readable storage medium that stores instructions that when executed by a processor cause the processor to process an image using a pre-trained three-dimension (3D) face generative adversarial network (GAN) having a generator and a neural renderer by per-forming the steps of: processing the image including a face; improving face latent code and camera parameters of the 3D face GAN by; freezing the generator and the neural renderer to improve face latent code and camera parameters; adjusting a face-to-camera distance and a camera focal length while maintaining a pupillary distance of the face constant; and rendering a photorealistic image; tuning a generator of the 3D face GAN; and generating an improved image of the face with reduced face distortion due to a longer camera-to-face distance. ∗ ∗ ∗ ∗ ∗
Device structures
Layer stacks claimed or described, ordered top of device to substrate.
pre-trained 3D face generative adversarial network (GAN)
discriminatordiscriminator
neuralrendererneural renderer
generatorgenerator
Reported properties
Performance values and ranges asserted in the specification or claims.
Property
Value
Material
Thickness
20–60 cm
—
Cited prior art
Patents and literature cited by this patent (applicant and examiner references).
Cited patents · 5
US 2009/0059041 A12009/0059041 A1 3/2009 Kwon
US 2022/0138455 A12022/0138455 A1 5/2022 Nagano et al.
US 2023/0245330 A12023/0245330 A1 * 8/2023 Nadir........................ G06T 5/60examiner
US 2023/0252714 A12023/0252714 A1 * 8/2023 Bradley................ G06T 15/205examiner
US 2024/0193891 A12024/0193891 A1 6/2024 Markhasin et al.
Efficient Geometry-aware 3D Generative Adversarial Networks (Year: 2022).
Perspective-Aware Manipulation of Portrait Photos. Fried, Ohad, et al., “Perspective-Aware Manipulation of Portrait Photos”, ACM Transactions on Graphics, vol. 35, Issue 4, Article 128, Jul. 11, 2016, pp. 1-10.
Learning Perspective Undistortion of Portraits. Zhao, Yajie, et al., “Learning Perspective Undistortion of Portraits”, IEEE/CVF International Conference on Computer Vision (ICCV),
Seoul, Korea, Oct. 27, 2019-Nov. 2, 2019, pp. 7849-7859.
FLARE: Fast learning of animatable and relightable mesh avatars. Bharadwaj, Shrisha et al. “FLARE: Fast learning of animatable and relightable mesh avatars”. arXiv preprint arXiv:2310.17519 (2023).
Efficient Geometry-aware 3D Generative Adversarial Networks. Chan et al: “Efficient Geometry-aware 3D Generative Adversarial Networks”, arxiv.org, Cornell University, Ithaca, NY, Apr. 27, 2022 (Apr. 27, 2022), XP091195806. Daneˇcˇek, Radek et al. “Emoca: Emotion driven monocular face capture and animation”. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (2022), pp. 20311-20322.
Arcface: Additive angular margin loss for deep face recognition. Deng, Jiankang et al. “Arcface: Additive angular margin loss for deep face recognition”. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. (2019a) pp. 4690-4699. Deng, Yu et al. “Accurate 3d face reconstruction with weakly- supervised learning: From single image to image set”. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition workshops. (2019b) 0-0.
Hugging Face. Face Parsing.. Dinu, Jonathan: “Hugging Face. Face Parsing.” (2021) https://huggingface.co/jonathandinu/face-parsing.arxiv:2105.15203 [Online; accessed Jan. 1, 2024].
Learning an animatable detailed 3D face model from in-the-wild images. Feng, Yao et al.: “Learning an animatable detailed 3D face model from in-the-wild images”. ACM Transactions on Graphics, (ToG) 40, 4 (2021), pp. 1-13.
Generative adversarial nets. Goodfellow, Ian et al.: “Generative adversarial nets”. Advances in neural information processing systems 27 (2014).
Improved training of wasserstein gans. Gulrajani, Ishaan et al.: “Improved training of wasserstein gans”. Advances in neural information processing systems 30 (2017).
Perspective reconstruction of human faces by joint mesh and landmark regression. Guo, Jia et al.: “Perspective reconstruction of human faces by joint mesh and landmark regression”. In European Conference on Com- puter Vision. (2022) Springer, 350-365. International Search Report and Written Opinion for International Application No. PCT/US2024/015996, dated Jun. 26, 2024 (Jun. 26, 2024)—14 pages.
Image-to-image translation with conditional adversarial networks. Isola, Phillip et al.: “Image-to-image translation with conditional adversarial networks”. In Proceedings of the IEEE conference on computer vision and pattern recognition. (2017) pp. 1125-1134.
Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image. Kao, Yueying et al.: “Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image”. IEEE Transactions on Image Processing (2023).
Realistic one-shot mesh-based head ava- tars. Khakhulin, Taras et al.: “Realistic one-shot mesh-based head ava- tars”. In European Conference on Computer Vision. (2022) Springer, pp. 345-362.
Adam: A method for stochastic opti- mization.. Kingma, Diederik P. et al.: “Adam: A method for stochastic opti- mization.” arXiv preprint arXiv:1412.6980 (2014).
Efficient multispectral facial capture with monochrome cameras.. Legendre, Chloe et al.: “Efficient multispectral facial capture with monochrome cameras.” In ACM SIGGRAPH 2018 Posters. pp. 1-2.
Learning a model of facial shape and expression from 4D scans.. Li, Tianye et al.: “Learning a model of facial shape and expression from 4D scans.” ACM Trans. Graph. 36, 6 (2017), 194-1.
3D GAN Inversion for Controllable Portrait Image Animation. Lin et al.: “3D GAN Inversion for Controllable Portrait Image Animation”, arxiv.org, Cornell University, Ithaca, NY, Mar. 25, 2022 (Mar. 25, 2022), XP091184354.
Mediapipe: A framework for building perception pipelines.. Lugaresi, Camillo et al.: “Mediapipe: A framework for building perception pipelines.” arXiv preprint arXiv:1906.08172 (2019). Nvidia. (2019). “Flickr-Faces-HQ Dataset (FFHQ)”. https://github. com/NVlabs/ffhq-dataset. [Online; accessed Jan. 1, 2024].
Image-to-image translation: Methods and applications.. Pang, Yingxue et al.: “Image-to-image translation: Methods and applications.” IEEE Transactions on Multimedia 24 (2021), pp. 3859-3881.
A 3D face model for pose and illumination invariant face recognition. Paysan, Pascal et al.: “A 3D face model for pose and illumination invariant face recognition”. In 2009 sixth IEEE international con- ference on advanced video and signal based surveillance. (2009) IEEE, pp. 296-301.
Multi-view 3D face reconstruction in the wild using siamese networks. Ramon, Eduard et al.: “Multi-view 3D face reconstruction in the wild using siamese networks”. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. (2019) 0-0.
U-net: Convolutional networks for bio- medical image segmentation. Ronneberger, Olaf et al.: “U-net: Convolutional networks for bio- medical image segmentation”. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015: 18th Interna- tional Conference, Munich, Germany, Oct. 5-9, 2015, Proceedings, Part III 18. Springer, pp. 234-241.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Ruiz, Nataniel et al.: “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation”. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec- ognition. (2023) pp. 22500-22510. Shih, YiChang et al.: “Distortion-free wide-angle portraits on cam- era phones”. ACM Transactions on Graphics (TOG) 38, 4 (2019), pp. 1-12.
Photo-Realistic 360° Head Avatars in the Wild. Szymanowicz, Stanislaw et al.: “Photo-Realistic 360° Head Avatars in the Wild”. In European Conference on Computer Vision. (2022) Springer, pp. 660-667.
Perspective distortion modeling, learning and compensation. Valente, Joachim et al.: “Perspective distortion modeling, learning and compensation”. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. (2015) pp. 9-16.
DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs. Wang et al.: “DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs”, arxiv.org, Cornell University, Ithaca, NY, Feb. 23, 2023 (Feb. 23, 2023), XP091449486.
3d face reconstruction with dense landmarks. Wood, Erroll et al.: “3d face reconstruction with dense landmarks”. In European Conference on Computer Vision. (2022) Springer, 160-177.
The unreasonable effectiveness of deep features as a perceptual metric. Zhang, Richard et al.: “The unreasonable effectiveness of deep features as a perceptual metric”. In Proceedings of the IEEE conference on computer vision and pattern recognition (2018), pp. 586-595. Zhu, Jun-Yan et al.: “Unpaired image-to-image translation using cycle-consistent adversarial networks”. In Proceedings of the IEEE international conference on computer vision. (2017) pp. 2223-2232.
Towards metrical reconstruction of human faces. Zielonka, Wojciech et al.: “Towards metrical reconstruction of human faces”. In European Conference on Computer Vision. (2022) Springer, pp. 250-269. International Search Report and Written Opinion for PCT/US2025/034980 dated Oct. 8, 2025, 12 pages.
Differentiable Rendering: A Survey. Kato, Hiroharu et al.: “Differentiable Rendering: A Survey,” Arxiv. org,CornellUniversityLibrary,Ithaca,NY,Jul.31,2020,XP081726202.
High-Quality Passive Facial Performance Capture Using Anchor Frames.. Beeler, Thabo et al. “High-Quality Passive Facial Performance Capture Using Anchor Frames.” Ed. by Hugues Hoppe. ACM SIGGRAPH 2011 papers 30.4 (2011): 1-10. Web.
Neural 3D Mesh Renderer.. Kato, Hiroharu et al.: “Neural 3D Mesh Renderer.” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018.
Efficient Geometry-aware 3D Generative Adversarial Networks (Year: 2022).
Perspective-Aware Manipulation of Portrait Photos. Fried, Ohad, et al., “Perspective-Aware Manipulation of Portrait Photos”, ACM Transactions on Graphics, vol. 35, Issue 4, Article 128, Jul. 11, 2016, pp. 1-10.
Learning Perspective Undistortion of Portraits. Zhao, Yajie, et al., “Learning Perspective Undistortion of Portraits”, IEEE/CVF International Conference on Computer Vision (ICCV),
Seoul, Korea, Oct. 27, 2019-Nov. 2, 2019, pp. 7849-7859.
FLARE: Fast learning of animatable and relightable mesh avatars. Bharadwaj, Shrisha et al. “FLARE: Fast learning of animatable and relightable mesh avatars”. arXiv preprint arXiv:2310.17519 (2023).
Efficient Geometry-aware 3D Generative Adversarial Networks. Chan et al: “Efficient Geometry-aware 3D Generative Adversarial Networks”, arxiv.org, Cornell University, Ithaca, NY, Apr. 27, 2022 (Apr. 27, 2022), XP091195806. Daneˇcˇek, Radek et al. “Emoca: Emotion driven monocular face capture and animation”. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (2022), pp. 20311-20322.
Arcface: Additive angular margin loss for deep face recognition. Deng, Jiankang et al. “Arcface: Additive angular margin loss for deep face recognition”. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. (2019a) pp. 4690-4699. Deng, Yu et al. “Accurate 3d face reconstruction with weakly- supervised learning: From single image to image set”. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition workshops. (2019b) 0-0.
Hugging Face. Face Parsing.. Dinu, Jonathan: “Hugging Face. Face Parsing.” (2021) https://huggingface.co/jonathandinu/face-parsing.arxiv:2105.15203 [Online; accessed Jan. 1, 2024].
Learning an animatable detailed 3D face model from in-the-wild images. Feng, Yao et al.: “Learning an animatable detailed 3D face model from in-the-wild images”. ACM Transactions on Graphics, (ToG) 40, 4 (2021), pp. 1-13.
Generative adversarial nets. Goodfellow, Ian et al.: “Generative adversarial nets”. Advances in neural information processing systems 27 (2014).
Improved training of wasserstein gans. Gulrajani, Ishaan et al.: “Improved training of wasserstein gans”. Advances in neural information processing systems 30 (2017).
Perspective reconstruction of human faces by joint mesh and landmark regression. Guo, Jia et al.: “Perspective reconstruction of human faces by joint mesh and landmark regression”. In European Conference on Com- puter Vision. (2022) Springer, 350-365. International Search Report and Written Opinion for International Application No. PCT/US2024/015996, dated Jun. 26, 2024 (Jun. 26, 2024)—14 pages.
Image-to-image translation with conditional adversarial networks. Isola, Phillip et al.: “Image-to-image translation with conditional adversarial networks”. In Proceedings of the IEEE conference on computer vision and pattern recognition. (2017) pp. 1125-1134.
Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image. Kao, Yueying et al.: “Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image”. IEEE Transactions on Image Processing (2023).
Realistic one-shot mesh-based head ava- tars. Khakhulin, Taras et al.: “Realistic one-shot mesh-based head ava- tars”. In European Conference on Computer Vision. (2022) Springer, pp. 345-362.
Adam: A method for stochastic opti- mization.. Kingma, Diederik P. et al.: “Adam: A method for stochastic opti- mization.” arXiv preprint arXiv:1412.6980 (2014).
Efficient multispectral facial capture with monochrome cameras.. Legendre, Chloe et al.: “Efficient multispectral facial capture with monochrome cameras.” In ACM SIGGRAPH 2018 Posters. pp. 1-2.
Learning a model of facial shape and expression from 4D scans.. Li, Tianye et al.: “Learning a model of facial shape and expression from 4D scans.” ACM Trans. Graph. 36, 6 (2017), 194-1.
3D GAN Inversion for Controllable Portrait Image Animation. Lin et al.: “3D GAN Inversion for Controllable Portrait Image Animation”, arxiv.org, Cornell University, Ithaca, NY, Mar. 25, 2022 (Mar. 25, 2022), XP091184354.
Mediapipe: A framework for building perception pipelines.. Lugaresi, Camillo et al.: “Mediapipe: A framework for building perception pipelines.” arXiv preprint arXiv:1906.08172 (2019). Nvidia. (2019). “Flickr-Faces-HQ Dataset (FFHQ)”. https://github. com/NVlabs/ffhq-dataset. [Online; accessed Jan. 1, 2024].
Image-to-image translation: Methods and applications.. Pang, Yingxue et al.: “Image-to-image translation: Methods and applications.” IEEE Transactions on Multimedia 24 (2021), pp. 3859-3881.
A 3D face model for pose and illumination invariant face recognition. Paysan, Pascal et al.: “A 3D face model for pose and illumination invariant face recognition”. In 2009 sixth IEEE international con- ference on advanced video and signal based surveillance. (2009) IEEE, pp. 296-301.
Multi-view 3D face reconstruction in the wild using siamese networks. Ramon, Eduard et al.: “Multi-view 3D face reconstruction in the wild using siamese networks”. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. (2019) 0-0.
U-net: Convolutional networks for bio- medical image segmentation. Ronneberger, Olaf et al.: “U-net: Convolutional networks for bio- medical image segmentation”. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015: 18th Interna- tional Conference, Munich, Germany, Oct. 5-9, 2015, Proceedings, Part III 18. Springer, pp. 234-241.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Ruiz, Nataniel et al.: “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation”. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec- ognition. (2023) pp. 22500-22510. Shih, YiChang et al.: “Distortion-free wide-angle portraits on cam- era phones”. ACM Transactions on Graphics (TOG) 38, 4 (2019), pp. 1-12.
Photo-Realistic 360° Head Avatars in the Wild. Szymanowicz, Stanislaw et al.: “Photo-Realistic 360° Head Avatars in the Wild”. In European Conference on Computer Vision. (2022) Springer, pp. 660-667.
Perspective distortion modeling, learning and compensation. Valente, Joachim et al.: “Perspective distortion modeling, learning and compensation”. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. (2015) pp. 9-16.
DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs. Wang et al.: “DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs”, arxiv.org, Cornell University, Ithaca, NY, Feb. 23, 2023 (Feb. 23, 2023), XP091449486.
3d face reconstruction with dense landmarks. Wood, Erroll et al.: “3d face reconstruction with dense landmarks”. In European Conference on Computer Vision. (2022) Springer, 160-177.
The unreasonable effectiveness of deep features as a perceptual metric. Zhang, Richard et al.: “The unreasonable effectiveness of deep features as a perceptual metric”. In Proceedings of the IEEE conference on computer vision and pattern recognition (2018), pp. 586-595. Zhu, Jun-Yan et al.: “Unpaired image-to-image translation using cycle-consistent adversarial networks”. In Proceedings of the IEEE international conference on computer vision. (2017) pp. 2223-2232.
Towards metrical reconstruction of human faces. Zielonka, Wojciech et al.: “Towards metrical reconstruction of human faces”. In European Conference on Computer Vision. (2022) Springer, pp. 250-269. International Search Report and Written Opinion for PCT/US2025/034980 dated Oct. 8, 2025, 12 pages.
Differentiable Rendering: A Survey. Kato, Hiroharu et al.: “Differentiable Rendering: A Survey,” Arxiv. org,CornellUniversityLibrary,Ithaca,NY,Jul.31,2020,XP081726202.
High-Quality Passive Facial Performance Capture Using Anchor Frames.. Beeler, Thabo et al. “High-Quality Passive Facial Performance Capture Using Anchor Frames.” Ed. by Hugues Hoppe. ACM SIGGRAPH 2011 papers 30.4 (2011): 1-10. Web.
Neural 3D Mesh Renderer.. Kato, Hiroharu et al.: “Neural 3D Mesh Renderer.” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018.
Efficient Geometry-aware 3D Generative Adversarial Networks (Year: 2022).
Perspective-Aware Manipulation of Portrait Photos. Fried, Ohad, et al., “Perspective-Aware Manipulation of Portrait Photos”, ACM Transactions on Graphics, vol. 35, Issue 4, Article 128, Jul. 11, 2016, pp. 1-10.
Learning Perspective Undistortion of Portraits. Zhao, Yajie, et al., “Learning Perspective Undistortion of Portraits”, IEEE/CVF International Conference on Computer Vision (ICCV),
Seoul, Korea, Oct. 27, 2019-Nov. 2, 2019, pp. 7849-7859.
FLARE: Fast learning of animatable and relightable mesh avatars. Bharadwaj, Shrisha et al. “FLARE: Fast learning of animatable and relightable mesh avatars”. arXiv preprint arXiv:2310.17519 (2023).
Efficient Geometry-aware 3D Generative Adversarial Networks. Chan et al: “Efficient Geometry-aware 3D Generative Adversarial Networks”, arxiv.org, Cornell University, Ithaca, NY, Apr. 27, 2022 (Apr. 27, 2022), XP091195806. Daneˇcˇek, Radek et al. “Emoca: Emotion driven monocular face capture and animation”. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (2022), pp. 20311-20322.
Arcface: Additive angular margin loss for deep face recognition. Deng, Jiankang et al. “Arcface: Additive angular margin loss for deep face recognition”. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. (2019a) pp. 4690-4699. Deng, Yu et al. “Accurate 3d face reconstruction with weakly- supervised learning: From single image to image set”. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition workshops. (2019b) 0-0.
Hugging Face. Face Parsing.. Dinu, Jonathan: “Hugging Face. Face Parsing.” (2021) https://huggingface.co/jonathandinu/face-parsing.arxiv:2105.15203 [Online; accessed Jan. 1, 2024].
Learning an animatable detailed 3D face model from in-the-wild images. Feng, Yao et al.: “Learning an animatable detailed 3D face model from in-the-wild images”. ACM Transactions on Graphics, (ToG) 40, 4 (2021), pp. 1-13.
Generative adversarial nets. Goodfellow, Ian et al.: “Generative adversarial nets”. Advances in neural information processing systems 27 (2014).
Improved training of wasserstein gans. Gulrajani, Ishaan et al.: “Improved training of wasserstein gans”. Advances in neural information processing systems 30 (2017).
Perspective reconstruction of human faces by joint mesh and landmark regression. Guo, Jia et al.: “Perspective reconstruction of human faces by joint mesh and landmark regression”. In European Conference on Com- puter Vision. (2022) Springer, 350-365. International Search Report and Written Opinion for International Application No. PCT/US2024/015996, dated Jun. 26, 2024 (Jun. 26, 2024)—14 pages.
Image-to-image translation with conditional adversarial networks. Isola, Phillip et al.: “Image-to-image translation with conditional adversarial networks”. In Proceedings of the IEEE conference on computer vision and pattern recognition. (2017) pp. 1125-1134.
Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image. Kao, Yueying et al.: “Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image”. IEEE Transactions on Image Processing (2023).
Realistic one-shot mesh-based head ava- tars. Khakhulin, Taras et al.: “Realistic one-shot mesh-based head ava- tars”. In European Conference on Computer Vision. (2022) Springer, pp. 345-362.
Adam: A method for stochastic opti- mization.. Kingma, Diederik P. et al.: “Adam: A method for stochastic opti- mization.” arXiv preprint arXiv:1412.6980 (2014).
Efficient multispectral facial capture with monochrome cameras.. Legendre, Chloe et al.: “Efficient multispectral facial capture with monochrome cameras.” In ACM SIGGRAPH 2018 Posters. pp. 1-2.
Learning a model of facial shape and expression from 4D scans.. Li, Tianye et al.: “Learning a model of facial shape and expression from 4D scans.” ACM Trans. Graph. 36, 6 (2017), 194-1.
3D GAN Inversion for Controllable Portrait Image Animation. Lin et al.: “3D GAN Inversion for Controllable Portrait Image Animation”, arxiv.org, Cornell University, Ithaca, NY, Mar. 25, 2022 (Mar. 25, 2022), XP091184354.
Mediapipe: A framework for building perception pipelines.. Lugaresi, Camillo et al.: “Mediapipe: A framework for building perception pipelines.” arXiv preprint arXiv:1906.08172 (2019). Nvidia. (2019). “Flickr-Faces-HQ Dataset (FFHQ)”. https://github. com/NVlabs/ffhq-dataset. [Online; accessed Jan. 1, 2024].
Image-to-image translation: Methods and applications.. Pang, Yingxue et al.: “Image-to-image translation: Methods and applications.” IEEE Transactions on Multimedia 24 (2021), pp. 3859-3881.
A 3D face model for pose and illumination invariant face recognition. Paysan, Pascal et al.: “A 3D face model for pose and illumination invariant face recognition”. In 2009 sixth IEEE international con- ference on advanced video and signal based surveillance. (2009) IEEE, pp. 296-301.
Multi-view 3D face reconstruction in the wild using siamese networks. Ramon, Eduard et al.: “Multi-view 3D face reconstruction in the wild using siamese networks”. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. (2019) 0-0.
U-net: Convolutional networks for bio- medical image segmentation. Ronneberger, Olaf et al.: “U-net: Convolutional networks for bio- medical image segmentation”. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015: 18th Interna- tional Conference, Munich, Germany, Oct. 5-9, 2015, Proceedings, Part III 18. Springer, pp. 234-241.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Ruiz, Nataniel et al.: “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation”. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec- ognition. (2023) pp. 22500-22510. Shih, YiChang et al.: “Distortion-free wide-angle portraits on cam- era phones”. ACM Transactions on Graphics (TOG) 38, 4 (2019), pp. 1-12.
Photo-Realistic 360° Head Avatars in the Wild. Szymanowicz, Stanislaw et al.: “Photo-Realistic 360° Head Avatars in the Wild”. In European Conference on Computer Vision. (2022) Springer, pp. 660-667.
Perspective distortion modeling, learning and compensation. Valente, Joachim et al.: “Perspective distortion modeling, learning and compensation”. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. (2015) pp. 9-16.
DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs. Wang et al.: “DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs”, arxiv.org, Cornell University, Ithaca, NY, Feb. 23, 2023 (Feb. 23, 2023), XP091449486.
3d face reconstruction with dense landmarks. Wood, Erroll et al.: “3d face reconstruction with dense landmarks”. In European Conference on Computer Vision. (2022) Springer, 160-177.
The unreasonable effectiveness of deep features as a perceptual metric. Zhang, Richard et al.: “The unreasonable effectiveness of deep features as a perceptual metric”. In Proceedings of the IEEE conference on computer vision and pattern recognition (2018), pp. 586-595. Zhu, Jun-Yan et al.: “Unpaired image-to-image translation using cycle-consistent adversarial networks”. In Proceedings of the IEEE international conference on computer vision. (2017) pp. 2223-2232.
Towards metrical reconstruction of human faces. Zielonka, Wojciech et al.: “Towards metrical reconstruction of human faces”. In European Conference on Computer Vision. (2022) Springer, pp. 250-269. International Search Report and Written Opinion for PCT/US2025/034980 dated Oct. 8, 2025, 12 pages.
Differentiable Rendering: A Survey. Kato, Hiroharu et al.: “Differentiable Rendering: A Survey,” Arxiv. org,CornellUniversityLibrary,Ithaca,NY,Jul.31,2020,XP081726202.
High-Quality Passive Facial Performance Capture Using Anchor Frames.. Beeler, Thabo et al. “High-Quality Passive Facial Performance Capture Using Anchor Frames.” Ed. by Hugues Hoppe. ACM SIGGRAPH 2011 papers 30.4 (2011): 1-10. Web.
Neural 3D Mesh Renderer.. Kato, Hiroharu et al.: “Neural 3D Mesh Renderer.” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018.
Efficient Geometry-aware 3D Generative Adversarial Networks (Year: 2022).
Perspective-Aware Manipulation of Portrait Photos. Fried, Ohad, et al., “Perspective-Aware Manipulation of Portrait Photos”, ACM Transactions on Graphics, vol. 35, Issue 4, Article 128, Jul. 11, 2016, pp. 1-10.
Learning Perspective Undistortion of Portraits. Zhao, Yajie, et al., “Learning Perspective Undistortion of Portraits”, IEEE/CVF International Conference on Computer Vision (ICCV),
Seoul, Korea, Oct. 27, 2019-Nov. 2, 2019, pp. 7849-7859.
FLARE: Fast learning of animatable and relightable mesh avatars. Bharadwaj, Shrisha et al. “FLARE: Fast learning of animatable and relightable mesh avatars”. arXiv preprint arXiv:2310.17519 (2023).
Efficient Geometry-aware 3D Generative Adversarial Networks. Chan et al: “Efficient Geometry-aware 3D Generative Adversarial Networks”, arxiv.org, Cornell University, Ithaca, NY, Apr. 27, 2022 (Apr. 27, 2022), XP091195806. Daneˇcˇek, Radek et al. “Emoca: Emotion driven monocular face capture and animation”. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (2022), pp. 20311-20322.
Arcface: Additive angular margin loss for deep face recognition. Deng, Jiankang et al. “Arcface: Additive angular margin loss for deep face recognition”. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. (2019a) pp. 4690-4699. Deng, Yu et al. “Accurate 3d face reconstruction with weakly- supervised learning: From single image to image set”. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition workshops. (2019b) 0-0.
Hugging Face. Face Parsing.. Dinu, Jonathan: “Hugging Face. Face Parsing.” (2021) https://huggingface.co/jonathandinu/face-parsing.arxiv:2105.15203 [Online; accessed Jan. 1, 2024].
Learning an animatable detailed 3D face model from in-the-wild images. Feng, Yao et al.: “Learning an animatable detailed 3D face model from in-the-wild images”. ACM Transactions on Graphics, (ToG) 40, 4 (2021), pp. 1-13.
Generative adversarial nets. Goodfellow, Ian et al.: “Generative adversarial nets”. Advances in neural information processing systems 27 (2014).
Improved training of wasserstein gans. Gulrajani, Ishaan et al.: “Improved training of wasserstein gans”. Advances in neural information processing systems 30 (2017).
Perspective reconstruction of human faces by joint mesh and landmark regression. Guo, Jia et al.: “Perspective reconstruction of human faces by joint mesh and landmark regression”. In European Conference on Com- puter Vision. (2022) Springer, 350-365. International Search Report and Written Opinion for International Application No. PCT/US2024/015996, dated Jun. 26, 2024 (Jun. 26, 2024)—14 pages.
Image-to-image translation with conditional adversarial networks. Isola, Phillip et al.: “Image-to-image translation with conditional adversarial networks”. In Proceedings of the IEEE conference on computer vision and pattern recognition. (2017) pp. 1125-1134.
Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image. Kao, Yueying et al.: “Towards 3d face reconstruction in perspective projection: Estimating 6dof face pose from monocular image”. IEEE Transactions on Image Processing (2023).
Realistic one-shot mesh-based head ava- tars. Khakhulin, Taras et al.: “Realistic one-shot mesh-based head ava- tars”. In European Conference on Computer Vision. (2022) Springer, pp. 345-362.
Adam: A method for stochastic opti- mization.. Kingma, Diederik P. et al.: “Adam: A method for stochastic opti- mization.” arXiv preprint arXiv:1412.6980 (2014).
Efficient multispectral facial capture with monochrome cameras.. Legendre, Chloe et al.: “Efficient multispectral facial capture with monochrome cameras.” In ACM SIGGRAPH 2018 Posters. pp. 1-2.
Learning a model of facial shape and expression from 4D scans.. Li, Tianye et al.: “Learning a model of facial shape and expression from 4D scans.” ACM Trans. Graph. 36, 6 (2017), 194-1.
3D GAN Inversion for Controllable Portrait Image Animation. Lin et al.: “3D GAN Inversion for Controllable Portrait Image Animation”, arxiv.org, Cornell University, Ithaca, NY, Mar. 25, 2022 (Mar. 25, 2022), XP091184354.
Mediapipe: A framework for building perception pipelines.. Lugaresi, Camillo et al.: “Mediapipe: A framework for building perception pipelines.” arXiv preprint arXiv:1906.08172 (2019). Nvidia. (2019). “Flickr-Faces-HQ Dataset (FFHQ)”. https://github. com/NVlabs/ffhq-dataset. [Online; accessed Jan. 1, 2024].
Image-to-image translation: Methods and applications.. Pang, Yingxue et al.: “Image-to-image translation: Methods and applications.” IEEE Transactions on Multimedia 24 (2021), pp. 3859-3881.
A 3D face model for pose and illumination invariant face recognition. Paysan, Pascal et al.: “A 3D face model for pose and illumination invariant face recognition”. In 2009 sixth IEEE international con- ference on advanced video and signal based surveillance. (2009) IEEE, pp. 296-301.
Multi-view 3D face reconstruction in the wild using siamese networks. Ramon, Eduard et al.: “Multi-view 3D face reconstruction in the wild using siamese networks”. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops. (2019) 0-0.
U-net: Convolutional networks for bio- medical image segmentation. Ronneberger, Olaf et al.: “U-net: Convolutional networks for bio- medical image segmentation”. In Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015: 18th Interna- tional Conference, Munich, Germany, Oct. 5-9, 2015, Proceedings, Part III 18. Springer, pp. 234-241.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Ruiz, Nataniel et al.: “Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation”. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec- ognition. (2023) pp. 22500-22510. Shih, YiChang et al.: “Distortion-free wide-angle portraits on cam- era phones”. ACM Transactions on Graphics (TOG) 38, 4 (2019), pp. 1-12.
Photo-Realistic 360° Head Avatars in the Wild. Szymanowicz, Stanislaw et al.: “Photo-Realistic 360° Head Avatars in the Wild”. In European Conference on Computer Vision. (2022) Springer, pp. 660-667.
Perspective distortion modeling, learning and compensation. Valente, Joachim et al.: “Perspective distortion modeling, learning and compensation”. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. (2015) pp. 9-16.
DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs. Wang et al.: “DisCO: Portrait Distortion Correction with Perspective- Aware 3D GANs”, arxiv.org, Cornell University, Ithaca, NY, Feb. 23, 2023 (Feb. 23, 2023), XP091449486.
3d face reconstruction with dense landmarks. Wood, Erroll et al.: “3d face reconstruction with dense landmarks”. In European Conference on Computer Vision. (2022) Springer, 160-177.
The unreasonable effectiveness of deep features as a perceptual metric. Zhang, Richard et al.: “The unreasonable effectiveness of deep features as a perceptual metric”. In Proceedings of the IEEE conference on computer vision and pattern recognition (2018), pp. 586-595. Zhu, Jun-Yan et al.: “Unpaired image-to-image translation using cycle-consistent adversarial networks”. In Proceedings of the IEEE international conference on computer vision. (2017) pp. 2223-2232.
Towards metrical reconstruction of human faces. Zielonka, Wojciech et al.: “Towards metrical reconstruction of human faces”. In European Conference on Computer Vision. (2022) Springer, pp. 250-269. International Search Report and Written Opinion for PCT/US2025/034980 dated Oct. 8, 2025, 12 pages.
Differentiable Rendering: A Survey. Kato, Hiroharu et al.: “Differentiable Rendering: A Survey,” Arxiv. org,CornellUniversityLibrary,Ithaca,NY,Jul.31,2020,XP081726202.
High-Quality Passive Facial Performance Capture Using Anchor Frames.. Beeler, Thabo et al. “High-Quality Passive Facial Performance Capture Using Anchor Frames.” Ed. by Hugues Hoppe. ACM SIGGRAPH 2011 papers 30.4 (2011): 1-10. Web.
Neural 3D Mesh Renderer.. Kato, Hiroharu et al.: “Neural 3D Mesh Renderer.” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE, 2018.