What Modern CryEngine Does To Your GPU | A Much Needed Revisit — Transcript
Full transcript
- 0:00I'm about to show you how a Kingdom Come
- 0:01Deliverance 2 frame is generated on your
- 0:03GPU by showing you how a frame is built
- 0:06draw category by draw category. This
- 0:08video is going to do three major things
- 0:10for you as long as you watch the whole
- 0:11video. You learning about optimized
- 0:13rendering is going to drive up market
- 0:15value for optimized content. By having a
- 0:18good idea of how the pipeline works,
- 0:19I'll be able to explain where even more
- 0:21performance can be driven by Warhorse
- 0:23Studios in future Cry Engine versions.
- 0:26Then by the end of the video, I will
- 0:28have shared with you how to enhance your
- 0:29native anti-alias experience without the
- 0:31need of a specific GPU brand or cost
- 0:33like you would find with FSR2. We're
- 0:36going to do the capture in this opening
- 0:37scene. With a full system restart, the
- 0:39game takes about 80% at most of the 12
- 0:42GB desktop 30 for a Vsync 60 fps target
- 0:45on settings just high enough to enable
- 0:47basic graphic features like
- 0:49long-distance foliage and local light
- 0:51shadow casting. For the capture, it's a
- 0:53forward jump motion to stress test any
- 0:55velocity related effects. Our captured
- 0:58frame also has a modified TA, which I'll
- 1:01explain how to enable when we reach the
- 1:02anti-aliasing segment in our analysis.
- 1:05This is the whole pipeline. And notice
- 1:07that the frame consists of 10,000 draws.
- 1:09And I would suggest referencing what
- 1:10we've seen in past analysis done on this
- 1:12channel because that is the highest
- 1:14we've seen. I'm going to switch out the
- 1:16metrics so our performance outliers can
- 1:17stand out a bit more. The first part of
- 1:19the pipeline are stencil clears and copy
- 1:21buffer regions totaling at around.32
- 1:23milliseconds. This is followed by a
- 1:25compute shader that processes cloud
- 1:26information at a 0.25 millisecond cost.
- 1:29The prepass processing is started but
- 1:31only for alpha tested foliage which is
- 1:33awful to see if you're familiar with the
- 1:35consequences of context processing.
- 1:37Luckily the prepass measures under a.3
- 1:39millisecond cost due to a small amount
- 1:40of draws and only utilizing a single
- 1:43opacity atlas texture reducing the cost
- 1:45of context switching. Removing the
- 1:47prepass in real time frees up around 3%
- 1:50of the GPU usage. I would take that 3%
- 1:52but unfortunately parts of the graphic
- 1:54pipeline are somehow dependent on the
- 1:56partial prepass which could be fixed if
- 1:58Warhorse Studios looked into it. The
- 2:01base pass is processed right afterwards
- 2:03writing to an albido with an unspecified
- 2:04alpha channel. A depth stencil which
- 2:07only has two channels. a world normal
- 2:09with an alpha containing a most likely
- 2:11specular, but it could be any of these
- 2:12three values, which I know four of which
- 2:14are written to the fourth render target.
- 2:16Now, for most of the analysis we've seen
- 2:18on the channel, most base pass object
- 2:20draws average around 15 to 25
- 2:22microconds. Most of the draws in this
- 2:25pipeline are 6 to 8 micro, which is 2 to
- 2:27three times faster than what we usually
- 2:29see. But there's an enormous amount of
- 2:31draws, and the outliers are in the high
- 2:3340s of microsconds, and some are
- 2:35tripledigit outliers.
- 2:37Let's look at one of the first base mass
- 2:38draws being a wooden structure. It's a
- 2:41good testament to the engine's draw
- 2:42sorting, but let's analyze the resources
- 2:44it took. For the amount of surface area
- 2:47it shades and the associated microcond
- 2:49reading, I would say this is a standard
- 2:51shading cost. But we can see that four
- 2:53textures were used for the material. And
- 2:55you want to use as little textures as
- 2:57possible in a material. You need
- 2:59textures to hold the color, normals,
- 3:01opacity, and other PBR elements so that
- 3:03the various shaders can transfer that to
- 3:05the GBuffers. Textures can only hold
- 3:07four channels of information. So, you
- 3:09can combine multiple PBR elements that
- 3:12are represented in a range from 1 to
- 3:14zero in a singular texture. When you
- 3:16have displacement textures, depth bias,
- 3:18ambient occlusion maps, you might be
- 3:20thinking you're running out of room and
- 3:21need another texture. Well, we have a
- 3:23lot of tricks. For instance, I've seen
- 3:25people combine specular cavity and
- 3:27ambient occlusion maps into a single
- 3:28channel and then use contrast math in
- 3:30the pixel shader to separate the values
- 3:32on the fly as they're written to the
- 3:34GBuffers. With normals, they usually
- 3:36take up three channels like base color
- 3:38information. But with normals, you can
- 3:41completely omit the Z channel and
- 3:42implement code in the shader that
- 3:44calculates Z with only X and Y values.
- 3:46These relatively simple on the-ly
- 3:49calculations are still cheaper than
- 3:50sampling a whole other texture. Many
- 3:52draws in Days Gone in this wood
- 3:54structure in Cry Engine show two channel
- 3:56normals which means that they're using
- 3:58this trick. But this draw here in KCD2
- 4:01has a separate texture that is only
- 4:02using one channel. So the opportunity to
- 4:04produce the total texture count isn't
- 4:06being taken here. Chances are if this is
- 4:09being done once, many other objects in
- 4:11the game are also failing to optimize in
- 4:13this way. Now, for this video, because
- 4:15this game has so many outliers among
- 4:17many, many small draws, I'm going to
- 4:19highlight the draws and then fade to the
- 4:21albido gbuffer so you can see what those
- 4:22draws are responsible for. Remember, if
- 4:25it's written to the albido, it's being
- 4:26written to the other buffers
- 4:28simultaneously. The way that I'll
- 4:29organize how many draws I highlight at a
- 4:31time is by recognizing the outliers are
- 4:33appearing. For a bit of context, outlier
- 4:37one is a textbook example of bloated
- 4:39texture count. Outliers two and three
- 4:41are just foliage draws. Outlier 4
- 4:44consists of teslated draws, which makes
- 4:46no sense considering their distance. If
- 4:49you have fast eyes, you can catch their
- 4:50frame contribution being largely
- 4:52overwritten by later draws. Notice that
- 4:55it's common for hundreds of draws to
- 4:57only make a small difference on the
- 4:58frame. It's also pretty clear that cheap
- 5:00terrain prepassing would have been
- 5:02extremely beneficial in preventing pixel
- 5:05overdraw. Then, objects that are skinned
- 5:07are rendered with a velocity buffer.
- 5:09These are the only objects with velocity
- 5:11information. With Silent Hill 2 and Jedi
- 5:14Survivor, velocity was rendered for
- 5:15every object as they were drawn to the
- 5:17base pass. The next portion of draws
- 5:20adds more detail to the terrain. And
- 5:22since terrain has been oluded by other
- 5:24objects by the maximum amount, you won't
- 5:26have to pay too much for the cost you
- 5:28would usually find with nonpass terrain
- 5:30rendering. The only issue here is that
- 5:32all these objects should have been
- 5:34prepassed instead of the foliage.
- 5:36Speaking of foliage, with a prepass
- 5:38adding a readed cost of 29 milliseconds,
- 5:40we get a total cost of 1.3 milliseconds
- 5:43for foliage. But let's talk about what
- 5:45assets are used to render the foliage.
- 5:47Most of the foliage uses the same
- 5:49texture atlas. The albido and opacity
- 5:51are used in the biggest texture, while
- 5:53the normal and separate indeterminate
- 5:55texture are drawn at half the
- 5:56resolution. The last texture is also
- 5:58indeterminate at an even lower
- 6:00resolution. As you can see, there is a
- 6:02missed opportunity to produce the total
- 6:03texture count once again. But let's talk
- 6:06about how the foliage is drawn to the
- 6:07frame. Analyzing the Zpass, we can see
- 6:09that nearby LODs don't incorporate a
- 6:12close cutout like Days Gone. They're
- 6:13just very simple rectangles. Now, more
- 6:16vertices are not only important for
- 6:17nearby LODs so that they aren't
- 6:19evaluating if opacity needs to be
- 6:21written in that area, but also for
- 6:23smoother animations. Considering the
- 6:25choppy look in motion and the fact that
- 6:26this foliage doesn't produce motion
- 6:28vectors, I thought we might be seeing a
- 6:29flipbook like approach for animation
- 6:32where the animation or motion that would
- 6:34be calculated in a vertex shader is
- 6:35actually stored as another set of
- 6:37vertices with different positions
- 6:39representing the motion. Vertex color is
- 6:41an easy way to identify sets of vertices
- 6:44within a singular mesh sets which you
- 6:46can specify to be called in a specific
- 6:48order of drawn frames. The lack of
- 6:50motion vectors could be due to no actual
- 6:53moving vertices and you end up with a
- 6:54finite range of motion or sets of
- 6:56vertices that represent that motion. You
- 6:59can see one of the geometric meshes used
- 7:01in one of the foliage draws that
- 7:02supports this idea. If you're a Cry
- 7:04Engine engineer, I suggest commenting on
- 7:06what the technique is more adjacent to.
- 7:08The next portion of the pipeline
- 7:10processes three 1600 by600 cascaded
- 7:13shadow maps with 5,000 individual draws.
- 7:16Most of them average around 5 to 6
- 7:17micro. The most expensive ones of course
- 7:20are the foliage draws. For a total cost
- 7:22for all shadow related passes, we total
- 7:24at 4.6 milliseconds. Mesh based depth is
- 7:28easy to render as long as quad overdraw
- 7:30isn't too bad like we see with these
- 7:31rocks. It takes around 1 millisecond to
- 7:34draw foliage shadows because referencing
- 7:36an opacity texture is far more
- 7:38complicated and that gets multiplied by
- 7:39the amount of surface area shaded in the
- 7:41shadow maps. Most games opt for screen
- 7:44space shadows and exclude foliage from
- 7:45the shadow pass. But what we see in this
- 7:47game is quality that could not be
- 7:49provided by a screen space technique.
- 7:51What we appreciate about this shot is
- 7:53presented in the first cascaded shadow
- 7:55map which just takes around.5
- 7:57milliseconds to render.4 milliseconds is
- 7:59just foliage. The next cascade takes
- 8:01another.5 milliseconds and renders way
- 8:04more foliage including foliage way
- 8:06behind the camera's frustm which is the
- 8:08screen space perspective. Grass uses
- 8:10instancing but many instances that will
- 8:12never cast shadows into the frustm are
- 8:14being rendered. The solution would be
- 8:15frusting with a lenient threshold. But
- 8:18the main problem here is that most of
- 8:20the foliage is so far away, meshes
- 8:22shouldn't be drawn anyway. It should at
- 8:24the very most use the open source Ben
- 8:27Studio approach and incorporate a shadow
- 8:29filling like distance field shadows,
- 8:31which are supposed to run faster than
- 8:32shadow maps anyway. Distance field
- 8:34shadows won't provide movement, but from
- 8:36this distance, it's not going to matter.
- 8:38You just want a little cheap fill to
- 8:40complement ambient and screen space
- 8:41techniques so you can retain a bit of
- 8:43that far silhouette. The third cascade
- 8:45is where hell breaks loose because it's
- 8:47drawing 5,000 individual trees, rocks,
- 8:50and a bunch of little objects and we can
- 8:52see the same thing happening in the base
- 8:53pass. There are serious problems with
- 8:56distant environment management and that
- 8:57may come down to a deficiency in HLOG
- 9:00generation in Cry Engine. These draws
- 9:02could be reduced with a few draws using
- 9:04instancing like we saw with the foliage
- 9:06draws or excluded like we saw in days
- 9:08gone. These rocks should all be one big
- 9:11lower poly mesh, not tiny rocks drawn
- 9:14one after another. This is also far
- 9:16enough to justify less precise distance
- 9:17field shadows as well. The next part is
- 9:20a copy resource draw followed by decal
- 9:22rendering which works without the
- 9:24prepass like frostbite 3 unlike UE4 to
- 9:275's overly restrictive pipeline.
- 9:29Processing Spogy related information in
- 9:31this scene costs 1.4 milliseconds and
- 9:33updates these quarter resolution 270p
- 9:35diffused lighting buffers which are then
- 9:37upscaled to native using the original
- 9:39depth information. This is followed by
- 9:41another draw that shades SSR data in a
- 9:43540p texture for a2 millisecond cost. 47
- 9:47micros are used to update this 540p SSR
- 9:49composition buffer. A pixel shader takes
- 9:5227 milliseconds to shade SSDO. Then
- 9:54another pixel shader is used to smooth
- 9:56out the noise while referencing the
- 9:58original depth for.1 millisecond cost.
- 10:00This is now the fastest AO I've analyzed
- 10:02on the channel, but I don't find it
- 10:04shading very convincing. The pipeline
- 10:06then creates a shadow mask with screen
- 10:08space shadows and the cascaded shadow
- 10:10maps. The cloud shadows are added for
- 10:11the biggest
- 10:12cost. 1.3 milliseconds is used to shade
- 10:15the lighting using almost all the
- 10:17G-buffers and lighting textures shown
- 10:18before this. 1 millisecond is used to
- 10:21shade subsurface gathered skin. The next
- 10:23portion is made out of a lot of little
- 10:24draws, but it downs samples a lot of
- 10:26shadows along with fog and cloud
- 10:28shading, taking about a millisecond to
- 10:29complete. Many instances of hair are
- 10:32forward rendered at the cost of 27
- 10:34milliseconds, but none of these draws
- 10:35indicate using MSAA. Cal's hair and Jedi
- 10:38Survivor has a1 millisecond reading
- 10:40total, and it was not drawn to the
- 10:42prepass. In this secondary capture, this
- 10:44character's hair takes4 milliseconds to
- 10:46render, but that's after a 029 microcond
- 10:49cost was induced to shade partial
- 10:50geometric information in the base pass.
- 10:53That's a 15 millisecond cost for hair. A
- 10:56copy resource takes.1 milliseconds and
- 10:58transparencies are rendered at the cost
- 10:59of 45 milliseconds. The next part is
- 11:02anti-aliasing. Three draws process the
- 11:04morphological aa SMA at the cost of.3
- 11:07milliseconds. Unfortunately, SMA tends
- 11:09to be butchered in many games including
- 11:11this one. Warframe has a pretty good
- 11:13implementation along with Crisis 2
- 11:15remastered when modifying it with a
- 11:16console. FXA is butchered often as well,
- 11:19but I'll talk about that another time.
- 11:21When SMA is done right, like many
- 11:23reshade implementations, it does a good
- 11:25job at removing jagged edges.
- 11:27Morphological anti-aliasing or MLA will
- 11:30never fix normal discontinuities or
- 11:33shader aliasing. Just like how non MLA
- 11:36based TIA with somewhat retainable image
- 11:38clarity will never solve jagged edges in
- 11:40motion, which is one of my major issues
- 11:42with this presentation from Epic Games.
- 11:45Nobody talks about this, but jagged
- 11:46edges always fall apart in motion when
- 11:49using TAA that has no morphological
- 11:51fallback. Everyone goes on and on about
- 11:53temporal stability when jitter TA edges
- 11:56act sporadically in motion. After SMA is
- 11:59applied to the image, the next draw
- 12:00takes 025 milliseconds to imprint the
- 12:02current frame on the accumulation frame.
- 12:05This is another capture with the SMA
- 12:07fallback applied in the same jumping
- 12:09forward motion. See that detail on the
- 12:11road? That's impressive considering how
- 12:13far away it is. Here's what happens when
- 12:15the stock temporal accumulation is
- 12:16added. Goodbye detail. We'll miss you.
- 12:19Take a look at the character outline.
- 12:21Stock temporal accumulation adds a
- 12:23trailing effect. Notice the texture
- 12:25destruction on this wooden fence. Look
- 12:27at this character ghosting all around
- 12:29him. And I'm sure plenty of people are
- 12:31familiar with the ghosting inside this
- 12:33game. Take a look at this last portion.
- 12:35Notice the cloth detail, the hair and
- 12:37armor detail, this person's facial hair
- 12:40and eyes. Notice the wood texture and
- 12:42the stand here. This other person's
- 12:44eyes, mouth, almost identifiable
- 12:46fingers, the person's face next to him.
- 12:48The dog's eyes and nose are visible. The
- 12:51texture on the rocks and pot all
- 12:53obliterated. Notice the aggressive
- 12:55averaging on trees. All of this clearly
- 12:58translates into aggressive blur. And
- 13:00that makes tons of sense when you
- 13:01realize the developers pumped up a
- 13:03temporal variable far from its less
- 13:05blurry default value. If you ask me,
- 13:07this is a bit too convenient for Nvidia
- 13:09when content creators publish comparison
- 13:11footage between native and DLSS quality.
- 13:14But even the default is too
- 13:15blurry/acumulative.
- 13:17If you modify the AA with these
- 13:19commands, you'll find that the first
- 13:20shot showing the detailed road now
- 13:22retains that detail. There is still an
- 13:25averaging behavior existing as you can
- 13:26see on the foliage and tree detail, but
- 13:28the texture detail and small detail
- 13:30features are retained. I would say your
- 13:33end results will look even better
- 13:34because this capture was using slightly
- 13:36different commands. Now, let's discuss
- 13:38this area right here. Kind of jagged,
- 13:40right? But this is why we know there's
- 13:42an issue with the morphological
- 13:43fallback. But even with those edges
- 13:45existing, you will not be able to
- 13:47perceive them because the next frame
- 13:48will jitter in a position where those
- 13:50edges will be filled. But only if you're
- 13:52over 59 frames per second with a VSYNC
- 13:54or maybe free sync alike option because
- 13:56those prevent frames from being skipped
- 13:58on your monitor, which is needed to
- 14:00prevent perceived jitter. If you're
- 14:02playing at a high frame rate, you can
- 14:03try increasing the pattern sequence. You
- 14:05have 15 options, but that doesn't
- 14:07correlate to the sequence count. For
- 14:09every 30 or so frames, you have the
- 14:11opportunity to increase the sequence
- 14:12length by one jitter position. To really
- 14:15stress my point to people who defend
- 14:17fake frames as the future of gaming,
- 14:19fake frames cannot provide resolution
- 14:21enhancement. This game provides DA and
- 14:23native FSR2. DA like FSR4 should have
- 14:27zero place in conversations because they
- 14:29are proprietary. If it's proprietary,
- 14:31it's worthless to me and it should be
- 14:33worthless to you. Proprietary features
- 14:35shouldn't even be mentioned by content
- 14:37creators. Now, you're going to have
- 14:39people say, "But FSR2 Native looks
- 14:41smoother than your modified
- 14:43SMA2TX." Yes, it does look smoother.
- 14:46FSR2 is also 10 times more expensive,
- 14:48and the smooth image it produces is an
- 14:51irrelevant lie because it's just built
- 14:53off of a relevant accumulation that
- 14:55falls apart in motion. If you feel like
- 14:58the modified SMA2TX isn't smooth enough,
- 15:00first of all, I'm not saying it's
- 15:01perfect, but it's very close and it just
- 15:03needs a few areas updated. We haven't
- 15:05even gone into additional AA methods
- 15:07like shader AA in MIT filtering. If you
- 15:10want to defend the temporary and
- 15:11irrelevant resolve you get from
- 15:12something like FSR2 during still motion,
- 15:15what you really want is high resolution
- 15:17gaming/ an expensive gaming setup. This
- 15:19channel combats the mass gaslighting
- 15:21existing within parts of the industry
- 15:23that push porta solutions as an even
- 15:25remotely similar experience to real
- 15:27high-end resolution gaming. The next
- 15:29portion is another copy resource that
- 15:31takes 0.1 milliseconds along with a
- 15:32pixel shader that copies another texture
- 15:34taking 75 microsconds. Motion blur
- 15:37processing takes.35 micros. Screen space
- 15:40bloom and sun shafts take2 milliseconds
- 15:42to process. 13 milliseconds is used for
- 15:45color grading while making a copy of
- 15:46that image takes 47 micro. Doof
- 15:49processing takes 4 milliseconds followed
- 15:52by a 1 millisecond call that shades a
- 15:53completely unnoticeable post aa grain
- 15:56effect. The next 38 microconds for the
- 15:58UI wraps up the pipeline at a 15.28
- 16:01microcond total timing. That's pretty
- 16:03accurate with a live timing, but more so
- 16:05than any other game we've analyzed, its
- 16:07pipeline efficiency, in other words, FPS
- 16:09will tank significantly if background
- 16:11applications such as an empty browser
- 16:13tab or a non-recording OBS window is
- 16:15opened on the system. Luckily,
- 16:16screenshots are fast enough to capture
- 16:18the more accurate median FPS timing,
- 16:20which matches our budget table. There
- 16:22are a number of conclusions this frame
- 16:24alone provides. the thoughts and
- 16:26potential regarding the pre-pass logic,
- 16:27how we should be processing hair and
- 16:29texture formats. You had a deeper look
- 16:31at anti-aliasing potential, how poor
- 16:33instancing consequences can trickle into
- 16:35multiple passes, poor LEDs on teslated
- 16:38objects, the importance of quality H law
- 16:40generation and shadow fallbacks like
- 16:42distance fields. You can now reference
- 16:44the efficiency of this foliage approach
- 16:46in terms of topology and texture
- 16:47management. I think we need more tests
- 16:49measuring Atlas performance, but that
- 16:51doesn't seem to be the most harmful
- 16:53aspect in terms of performance in
- 16:54comparison with the poor texture packing
- 16:57and the fact the GPU keeps switching
- 16:59memory context on such a chunky VRM
- 17:01asset. I think a major feature engines
- 17:04need to start offering artists is the
- 17:05ability to specify the order of any
- 17:07draws accessing an atlas, such as the
- 17:09end of the base pass before pre-passed
- 17:11objects are shaded. This is all part of
- 17:13a draw ordering approach our videos have
- 17:15outlined using real scenarios and engine
- 17:18producers need to start building on the
- 17:19clear logic we're showing. But I really
- 17:21want to take a quick moment to talk
- 17:23about the LOD transitions. Cry engine
- 17:25uses order dithering which is good but
- 17:27the transitions are tied to the camera.
- 17:30Not only does this allow for slow
- 17:32noticeable transitions that can stay
- 17:33midway visible if you stop moving but
- 17:35it's also allowing overdraw to last
- 17:37longer on screen which can build up its
- 17:39presence in the frame surface area. If
- 17:41you don't know what overdraw is,
- 17:43reference my Nanite video in the pin
- 17:44comment on it. Another big problem with
- 17:46the order dithering is that the pixel
- 17:48alignment is constant and not slightly
- 17:50jittered like we saw with some days gone
- 17:52objects. Cry engine fades between the
- 17:54full order dithered spectrum. These last
- 17:57two are pivotal in preventing the eye in
- 17:59noticing patterns. The last issue with
- 18:01the dithering is how shadows relate to
- 18:03the LOD transitions. The shadow of the
- 18:05newest LOD pops in immediately before
- 18:08the newest LOD fades into the main view.
- 18:11It should be a synchronized dithered
- 18:13transition, similar to what you see in
- 18:14my Nanite video at minute 920. But
- 18:17something I'd like to bring attention to
- 18:18is the fact that this channel has shown
- 18:20a lot of footage conveying the best
- 18:21possible way to transition LODs with
- 18:24dithering. But we do dithering because
- 18:26real fading requires translucent
- 18:28shading, which is expensive and not even
- 18:30possible with the fur de GBuffers. But
- 18:32ninth gen adjacent hardware might
- 18:34actually provide enough power to support
- 18:36real fading by temporarily forward
- 18:38rendering the LODs during transitions.
- 18:41The shadow maps would still need
- 18:42dithering that matches the forward
- 18:44rendered fade speed, but it's still
- 18:46interesting idea for current hardware.
- 18:48Now, let's discuss the lighting. I'm not
- 18:50a fan of SSDO because of its artifacting
- 18:53resolve. It is set to compute at a half
- 18:55resolution by default, but even when you
- 18:57switch to full resolution, you still end
- 18:59up with the same artifacts. It is very
- 19:01cheap to compute, so if Cry cared, they
- 19:03could work on making it a bit smoother
- 19:05for a little bit of an extra cost. But
- 19:07I'll be removing it from the pipeline in
- 19:09the rest of the footage. Now, Fogge
- 19:11caters to the crisis scenario, which is
- 19:12founded on destruction, which means it
- 19:14iterates over information. This game
- 19:16will never change. It's the same problem
- 19:18Fortnite introduces to Unreal. But even
- 19:21with these foundation similarities,
- 19:23Svogi's runtime cost is extremely
- 19:25stable, always remaining under two
- 19:27milliseconds, even in the most complex
- 19:28scenarios I tested it against. But there
- 19:31are important things to take notes on.
- 19:33Here is a scene with no GI and only
- 19:35direct lighting. Fogi has an obvious
- 19:37impact on visuals, but a few aspects are
- 19:39a bit simple. For instance, most of the
- 19:41additional indirect light comes from a
- 19:43skylight. A quick refresher on what a
- 19:45skylight does is light the entire scene
- 19:47in a uniform color, usually blue, to
- 19:50create ambient lighting. Skylighting is
- 19:52old, basically free, and people really
- 19:54need to start recognizing it because
- 19:55whenever RTGI is referred to as
- 19:57groundbreaking, it's always compared
- 19:59against pure skylighting combined with
- 20:01some disgustingly outdated
- 20:03SSAO. Like many somewhat competently lit
- 20:06games, this one combines local GI data
- 20:09with a skylight. Though the skylight
- 20:10value is pumped up at such a high
- 20:12amount, it drowns out a lot of the
- 20:14realistic aspects of the bounce light
- 20:16quality. The local GI provides large
- 20:18scale indirect shadowing along with
- 20:20local color bleed. I have doubts
- 20:22regarding the efficiency in terms of
- 20:24color balance and the difference between
- 20:25large scale AO offered by Unreal's DFAO.
- 20:28Sogi is also prone to temporal smearing
- 20:30like Lumen due to temporally rendering
- 20:32at a quarter% screen resolution. Like
- 20:35SSDO, I think the fast performance
- 20:37points towards a clear direction for
- 20:38visual upgrades and stability. Our next
- 20:41conclusion topic regards shadows. Other
- 20:43than the fact that I had to pump up the
- 20:45shadow resolution using the high shadow
- 20:47preset just to get local shadow casting,
- 20:49that of which enables at a comedic
- 20:51distance of pointlessness, the fact that
- 20:52it fades instead of popping on and off
- 20:55is far better than what we get with
- 20:56Unreal. Once again, let's remember that
- 20:58Unreal is one of the few engines that
- 21:00doesn't support this basic aesthetic
- 21:02optimization. The last thing I want to
- 21:04bring attention to is the texture and
- 21:06PBR workflow. For a lot of PBR
- 21:08scenarios, a lot of productions skip out
- 21:10on texture represented aspects and
- 21:11replace them with constant values for
- 21:13things like specular or completely omit
- 21:15pre-calculated aspects like AO maps.
- 21:18These four things are hindering realism
- 21:19in game materials, channel packing
- 21:22concerns, improper material evaluation
- 21:25caused by testing these texture maps
- 21:26with skylighting instead of HDR
- 21:28imagebased lighting. lack of in-gend
- 21:30workflows that allow rust or
- 21:32inexperienced productions to feed their
- 21:33texture/ channel values into
- 21:35standardized unpacking approaches. We're
- 21:38talking about untapped channel packing
- 21:40potential that could bring most object
- 21:41draws down to just two textures or three
- 21:44textures with a high amount of
- 21:45precomputed photorealistic properties.
- 21:47But in order to have good workflows, you
- 21:49have to have good defined standards in
- 21:51the shaders first. But that's a huge
- 21:54topic that this channel will have to
- 21:55cover in a specialized video. That wraps
- 21:57up this analysis. We hope you found it
- 21:59interesting and educational. If you did,
- 22:01press like and subscribe. And if you
- 22:03really want to support this kind of
- 22:04content and raise consumer awareness,
- 22:06share this video aggressively,
- 22:08specifically in Discord channels, gaming
- 22:10subreddits, forums, and repost your
- 22:12thoughts on social media. It takes about
- 22:14four weeks for us to produce new
- 22:16content. We have had to deal with
- 22:17various Pinocchios out there telling
- 22:19blatant lies and slander about our
- 22:21company. One person even made a claim
- 22:23that we've been fundraising money for
- 22:25over a year. when it's 100% provable,
- 22:27and you can check this for yourself,
- 22:28that we've only been publishing videos
- 22:30for about 9 months. Because many of you
- 22:32asked us how you can help, we gave
- 22:34simple instructions on our website on
- 22:36how to give our channel a super thanks.
- 22:38But after we did that, now we have
- 22:39supporters asking us to get a Patreon,
- 22:41and we are happy to announce we finally
- 22:43have. We had two missions when we began
- 22:45this YouTube journey. The first was to
- 22:47facilitate positive and powerful change
- 22:49in the consumer and developers. The
- 22:51second was to publish our game once UE5
- 22:54solutions were found. We have never
- 22:55wavered from our goals. Your support in
- 22:57the form of likes, subs, and super
- 22:59thanks help us grow and keep us going.
- 23:02Thank you so much to every single viewer
- 23:03and for your positive feedback.
About this transcript
This page contains the full transcript of What Modern CryEngine Does To Your GPU | A Much Needed Revisit by Threat Interactive, generated from the public captions YouTube serves with the video. The transcript has 4,422 words across 681 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.