KonstantinSeurer/mesa

Commit Graph

Author	SHA1	Message	Date
Kenneth Graunke	4f586cd8f1	i965: Push UBO data, but don't use it just yet. This patch starts uploading UBO data via 3DSTATE_CONSTANT_* packets, and updates the compiler to know that there's extra payload data, so things continue working. However, it still issues pull loads for all data. I wanted to separate the two aspects for greater bisectability. v2: Update for new intel_bufferobj_buffer parameter. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 20:18:30 -07:00
Kenneth Graunke	6834b1ebe3	i965: Pad buffer objects by 2kB in robust contexts to avoid OOB access. This is an annoyingly big hammer, but it seems less mean than disabling UBO pushing, and I'm not sure what else to do. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	67fec96452	i965: Stop re-uploading push constants after URB reconfiguration. Previously we would re-upload the constant data to the batchbuffer, then re-emit the packets. We only need to do the last step (causing the existing data in the batchbuffer to be re-uploaded to the push constant staging area in the L3). Now that we've separated the two, it's pretty easy to accomplish. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	29603ae208	i965: Separate uploading push constant data from the pointer packets. I hope to upload UBO via 3DSTATE_CONSTANT_XS packets, in addition to normal uniforms. In order to do that, I'll need to re-emit the packets when UBOs change. But I don't want to re-copy the regular uniform data to the batchbuffer every time. This patch separates out the data uploading from the packet submission. We're running low on dirty bits, so I made the new atom happen on every draw call, and added a flag to stage_state indicating that we want the packet for that stage emitted. I would have preferred to do this outside the atom system, but it has to happen between the uploading of push constant data and the binding table upload. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	a18ab92d3c	i965: Introduce a BRW_NEW_DRAW_CALL dirty bit. This allows us to have atoms which are signalled on every draw call. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	24891d7c05	i965: Store per-stage push constant BO pointers. Right now, we always upload new push constant data, and immediately emit 3DSTATE_CONSTANT_* packets. We call intel_upload_space and store the resulting BO pointer in brw->curbe.curbe_bo. We read that when emitting the packets. This works today, but is fragile - it depends on upload and packet emission being interleaved. If we instead were to upload all the data, then emit all the packets, then upload BO wrapping will get us into trouble. For example, the VS constants may land in one upload BO, but the FS constants may not fit and land in a second upload BO. Uploading FS constants would overwrite the brw->curbe.curbe_bo pointer, so when we emitted 3DSTATE_CONSTANT_VS, we'd get the wrong BO. I intend to separate out this code in a future commit, so I need to fix this. To fix it, we simply store a per-stage BO pointer. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	6d28c6e52c	i965: Select ranges of UBO data to be uploaded as push constants. This adds a NIR pass that decides which portions of UBOS we should upload as push constants, rather than pull constants. v2: Switch to uint16_t for the UBO block number, because we may have a lot of them in Vulkan (suggested by Jason). Add more comments about bitfield trickery (requested by Matt). v3: Skip vec4 stages for now...I haven't finished wiring up support in the vec4 backend, and so pushing the data but not using it will just be wasteful. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	2a5e4f15ef	i965: Require a UBO offset alignment of 32 bytes. Soon, we're going to start providing UBO data to shaders as push constants, rather than requiring them to issue pull loads. The 3DSTATE_CONSTANT_* commands require 32 byte aligned pointers. So, we need to increase this from 16 to 32. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	8ec5a4e4a4	i965: Switch to absolute addressing for constant buffer 0. By default, 3DSTATE_CONSTANT_* Constant Buffer 0 is relative to dynamic state base address. This makes it unusable for pushing UBOs. I'd like to be able to use all four push buffers. There is a bit in the INSTPM register (or CS_DEBUG_MODE2 on Skylake) which controls whether buffer 0 is relative to dynamic state base address, or simply a normal pointer. Setting that gives us full flexibility. We can't currently write this on Haswell and earlier, and will need to update the kernel command parser, and then do the whole version checking song and dance. Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-07-13 19:56:49 -07:00
Kenneth Graunke	86bd3fd864	i965: Use async maps for BufferSubData to regions with no valid data. When writing a region of a buffer via glBufferSubData(), we can write the data asynchronously if the destination doesn't contain any data. Even if it's busy, the data was undefined, so the new data is fine too. Removes all stall avoidance blits on BufferSubData calls in "Total War: WARHAMMER" on my Skylake GT4. Decreases the number of stall avoidance blits in Manhattan 3.1: - Skylake GT4: -18.3544% +/- 6.76483% (n=13) - Apollolake: -12.1095% +/- 5.24458% (n=13) Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2017-07-13 16:58:17 -07:00
Kenneth Graunke	5f223648f2	i965: Track a range of the buffer which contains valid data. Reviewed-by: Chris Wilson <chris@chris-wilson.co.uk> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2017-07-13 16:58:17 -07:00
Kenneth Graunke	f47612dafb	i965: Add a "write" parameter to intel_bufferobj_buffer. This doesn't do anything yet, but soon we'll want to know whether an access to a buffer section may write that data, or simply reads it. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2017-07-13 16:58:17 -07:00
Rafael Antognolli	9a9c7e452b	i965: Convert GS_STATE to genxml. Merge the code with gen6+ 3DSTATE_GS, and delete brw_gs_state.c, together with brw_gs_unit_state. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	2936388205	i965: Prepare gs_state emitting code to include gen4-5. Since we always call brw_batch_emit anyways, we can hopefully make things simpler by calling it only once, and then branching inside its body. This can be helpful when bringing the gen4-5 code into this function. Additionally, check for GEN_GEN == 6 instead of < 7 in cases that won't apply to lower gens. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	9a2cca929f	i965: Remove upload_gs_state_for_tf. This function only emits a particular case of 3DSTATE_GS. Instead, we can do that inside genX(upload_gs_state), and later reuse part of that code for emitting gen4-5 state. There's the additional benefit of allowing us to remove gen6_gs_state.c, which was only left because of this function. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	ad7663b838	i965: Convert BLEND_CONSTANT_COLOR state to genxml. It's a very simple conversion, and it allows us to delete brw_cc.c. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	5d48710981	i965: Convert CC state on gen4-5 to genxml. Use set_blend_entry_bits and set_depth_stencil_bits to fill most of the color calc struct, and then manually update the rest. v2: - Always check for depth_irb (Ken) - Always set Backface Stencil Ref (Ken) - Always set alpha reference value (Ken) Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	b0052aa46f	i965: Move color calc code around a bit. This makes the code more consistent accross generations. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	5246d3eb83	i965: Check for alpha channel just like in gen6+. gen6+ uses _mesa_base_format_has_channel() to check for the alpha channel, while gen4-5 use ctx->DrawBuffer->Visual.alphaBits. By using _mesa_base_format_has_channel() here we keep the same behavior accross all gen. While initially both ways of checking the alpha channel seemed correct to me, this change also seems to fix fbo-blending-formats piglit test on gen4. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Rafael Antognolli	1d2d3dbc8a	i965: Make a helper function for blend entry related state. Add a helper function to reuse code that fills blend entry related state, and make genX(upload_blend_state) use it. This function can later be used by gen4-5 color calc state to set the blend related bits. Signed-off-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 15:39:49 -07:00
Kenneth Graunke	e84cb56f48	i965: Make a helper function for depth/stencil related state. Gen4-5 basically glue DEPTH_STENCIL_STATE, COLOR_CALC_STATE, and BLEND_STATE together into a single COLOR_CALC_STATE structure. By making a helper function, we'll be able to reuse it when filling out Gen4-5 COLOR_CALC_STATE without replicating any actual logic. We use generation-defined typedef to handle the polymorphism. Reviewed-by: Rafael Antognolli <rafael.antognolli@intel.com>	2017-07-13 15:39:49 -07:00
Lionel Landwerlin	6131a1ae40	aubinator: don't leak fd of opened aubfile CID: 1373563 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:50 +01:00
Lionel Landwerlin	d1bd731e30	anv: don't use strcpy for copying strings CID: 1358935 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:47 +01:00
Lionel Landwerlin	226fae7849	intel/compiler: no need to check unsigned is >= 0 CID: 1338342 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:45 +01:00
Lionel Landwerlin	7c4daf8c37	i965: fix missing NULL return if allocation fails CID: 1250585 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:41 +01:00
Lionel Landwerlin	95c917668c	intel/compiler: don't check unsigned is >= 0 CID: 1224468 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:38 +01:00
Lionel Landwerlin	fd8e8fdbfe	i965: check pointer before dereferencing it Check that irb isn't NULL before accessing irb->Base.Base.NumSamples. CID: 1026046 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:35 +01:00
Lionel Landwerlin	b02d136b5e	i965: map_gtt: check mapping address before adding offset The NULL check might fail if offset isn't 0. CID: 971379 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:32 +01:00
Lionel Landwerlin	a25a533458	intel/compiler: remove check unsigned is >= 0 By definition unsigned are always >= 0. CID: 742212 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:29 +01:00
Lionel Landwerlin	19869d6091	isl: use 64bit arithmetic to compute size If we allow the size to be more than 2^32, then we should compute it in 64bit arithmetic otherwise we might run into overflow issues. CID: 1412892, 1412891 Signed-off-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-07-13 22:50:26 +01:00
Connor Abbott	4df93a54f1	nir/lower_io_to_temporaries: don't set compact on shadow vars The compact flag doesn't make sense on local variables, since the packing on them is up to the driver. This fixes nir_validate assertions in some cases, particularly when lower_io_to_temporaries is used on per-vertex inputs/outputs. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2017-07-13 14:45:25 -07:00
Connor Abbott	99ff7a9f1f	nir: don't segfault when printing variables with no name While normally we give variables whose name field is NULL a temporary name when called from nir_print_shader(), when we were calling from nir_print_instr() we never bothered, meaning that we just segfaulted when trying to print out instructions with such a variable. Since nir_print_instr() is meant to be called while debugging, we don't need to bother too much about giving a consistent name, but we don't want to crash in the middle of debugging. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl>	2017-07-13 14:40:23 -07:00
Jason Ekstrand	add72599d9	i965/urb: Trigger upload_urb on NEW_BLORP It's a bit rare, but blorp can trigger a urb reconfiguration. When that happens, we need to re-upload the URB config. Previoulsy blorp would set BRW_NEW_URB_SIZE, but this is a pretty big hammer as it would cause back-to-black blorp operations to reconfigure both times. Using BRW_NEW_BLORP is a small, more accurate hammer. v2 (idr): Sort BRW_NEW_ tokens to match brw_recalculate_urb_fence and gen6_urb. v3 (idr): Don't whack BRW_NEW_URB_SIZE in blorp. Suggested by Jason. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-07-13 13:41:09 -07:00
Kenneth Graunke	42c64b5f87	mesa: Return GL_INVALID_ENUM for bogus TEXTURE_SRGB_DECODE_EXT params. Fixes dEQP-GLES31.functional.debug.negative_coverage.get_error.shader.srgb_decode_samplerparameter{f,fv,i,Iiv,Iuiv,iv}. Reviewed-by: Jason Ekstrand <jason@jlekstrand.net> Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>	2017-07-13 13:00:58 -07:00
Marek Olšák	f33d8af7aa	st/dri: add 32-bit RGBX/RGBA formats Add support for 32-bit RGBX/RGBA formats which are required for Android. The original patch (commit `ccdcf91104`) was reverted (commit `c0c6ca40a2`) in mesa as it broke GLX resulting in swapped colors. Based on further investigation by Chad Versace, moving the RGBX/RGBA configs to the end is enough to prevent breaking GLX. The handling of RGBA/RGBX in dri_fill_st_visual is a fix from Marek Olšák. Cc: Eric Anholt <eric@anholt.net> Cc: Mauro Rossi <issor.oruam@gmail.com> Reviewed-by: Chad Versace <chadversary@chromium.org> Reviewed-by: Marek Olšák <marek.olsak@amd.com> Signed-off-by: Rob Herring <robh@kernel.org>	2017-07-13 14:36:47 -05:00
Eric Anholt	5a9fb2eabc	broadcom/vc4: Add more packets to the v2.1 XML. These will be used to replace vc4_cl_dump.c's hand-written dumping.	2017-07-13 11:30:42 -07:00
Eric Anholt	427bbbb99c	broadcom: Introduce a header for talking about chip revisions. This will be used by the VC5 driver and various shared VC4/VC5 tooling, like the XML decoder.	2017-07-13 11:28:28 -07:00
Eric Anholt	fd37ce6bec	broadcom/genxml: Use the same "gen" attr for HW version as Intel does. This will let us reuse their tools more easily.	2017-07-13 11:28:28 -07:00
Eric Anholt	ee170c9d83	broadcom/genxml: Support unpacking fixed-point fractional values. This was an oversight in the original XML support, because unpacking wasn't used much. The new XML-based CL dumper will want it, though.	2017-07-13 11:28:28 -07:00
Michel Dänzer	655a32f729	st/mesa: Handle st_framebuffer_create returning NULL st_framebuffer_create returns NULL if stfbi == NULL or st_framebuffer_add_renderbuffer returns false for the colour buffer. Fixes Xorg crashing on startup using glamor on radeonsi. Fixes: `147d7fb772` ("st/mesa: add a winsys buffers list in st_context") Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=101775 Signed-off-by: Michel Dänzer <michel.daenzer@amd.com> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Brian Paul <brianp@vmware.com>	2017-07-13 09:26:20 -06:00
Tim Rowley	254fa3dbf5	swr/rast: Fix use of KNL-only intrinsics in SKX build Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-13 08:47:10 -05:00
Tim Rowley	4c185dd3b3	swr/rast: Fix build warnings when using the Intel compiler Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-13 08:47:10 -05:00
Tim Rowley	bbc3b5c0dc	swr/rast: SIMD16 Frontend - Fix USE_SIMD16_FRONTEND build Previous check-ins without testing with USE_SIMD16_FRONTEND have introduced regressions. This fixes the build, not the regressions. Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-13 08:47:10 -05:00
Tim Rowley	640ea4d9a1	swr/rast: Removing unneeded MSVC warning pragma Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-13 08:47:10 -05:00
Tim Rowley	185b37f641	swr/rast: Add support for read-only render targets Core will ensure hot tiles are loaded for read and write render targets, and will skip all output merger for read-only render targets. Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-13 08:47:10 -05:00
Tim Rowley	d8ebcad540	swr/rast: Support render target mask instead of render target count WIP to support read-only render targets. Reviewed-by: Bruce Cherniak <bruce.cherniak@intel.com>	2017-07-13 08:47:10 -05:00
Alejandro Piñeiro	57671025b0	egl: remove unused err variable Fixes: `81e95924ea` ("egl: call _eglError within _eglParseImageAttribList") Reviewed-by: Daniel Stone <daniels@collabora.com>	2017-07-13 13:36:37 +02:00
Nicolai Hähnle	c22e3c5373	radeonsi/gfx9: fix crash building monolithic merged ES-GS shader Forwarding from the ES prolog to the ES just barely exceeds the current maximum array size when 16 vertex attributes are used. Give it a decent bump to account for merged shaders having up to 32 user SGPRs. Fixes a crash in GL45-CTS.multi_bind.draw_bind_vertex_buffers. Cc: mesa-stable@lists.freedesktop.org Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-07-13 13:01:15 +02:00
Thomas Hellstrom	81fb154777	loader/dri3: Use dri3_find_back in loader_dri3_swap_buffers_msc If the application hasn't done any drawing since the last call, we would reuse the same back buffer which was used for the previous swap, which may not have completed yet. This could result in various issues such as tearing or application hangs. In the normal case, the behaviour is unchanged. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=97957 Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=101683 Cc: mesa-stable@lists.freedesktop.org [Michel Dänzer: Make Thomas' fix from bugzilla actually work as intended, write commit log]	2017-07-13 16:49:28 +09:00
Jason Ekstrand	c3b5c2ca19	i965/screen: Drop get_tiled_height It's no longer used. Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-07-12 21:15:46 -07:00

1 2 3 4 5 ...

94003 Commits All Branches Search

94003 Commits

All Branches